fix(deepagents): gate cache_control writes on per-call request.model (#551)
Fixes #550.
## Problem
`createCacheBreakpointMiddleware` and `createMemoryMiddleware` write
Anthropic-specific `cache_control: { type: "ephemeral" }` markers into
the request payload. Both gate the write at **agent-creation time** in
`createDeepAgent` (whether the *primary* model is Anthropic). When
`modelFallbackMiddleware` swaps `request.model` to a non-Anthropic
provider at request time, the markers leak through and the fallback
provider rejects the request:
```
400 Unknown parameter: 'messages[0].content[N].cache_control'
```
Full root-cause writeup, production trace, and reproduction in #550.
## Fix
Move the model classification from boot-time to **per-call**, using the
existing `isAnthropicModel` helper (`src/utils.ts:9`).
- **`src/middleware/cache.ts`** — bail at the top of `wrapModelCall`
when `request.model` isn't Anthropic.
- **`src/middleware/memory.ts`** — AND the existing `addCacheControl`
flag with the same per-call check. `addCacheControl` keeps its current
public meaning (opt-in switch); the per-call gate is an additional
safety net.
- **`src/agent.ts`** — left untouched. The boot-time install gate stays
(small perf win when primary is non-Anthropic — middleware isn't
installed at all). The new per-call check handles the fallback case
where the middleware *is* installed but the model has been swapped.
## Tests
Each middleware's `.test.ts` got a `per-call provider gating` block.
They build a request with a `ChatAnthropic`-shaped stub vs a
`ChatOpenAI`-shaped stub on `request.model`, run `wrapModelCall`, and
assert `cache_control` is/isn't written. Also covers `ConfigurableModel`
with `modelProvider: "anthropic"` and the string forms
(`"anthropic:claude-..."`, `"openai:gpt-5"`).
Existing fixtures that didn't set `request.model` now pass a
`ChatAnthropic` stub so they keep exercising the write path — they were
previously relying on the unconditional behavior. No behavioral
assertions changed; only the fixture got more explicit.
31/31 tests pass on `libs/deepagents` (`cache.test.ts` +
`memory.test.ts`).
## Compatibility
- **No public API change.** `addCacheControl?: boolean` on
`MemoryMiddlewareOptions` keeps its existing semantics.
- **Anthropic-primary happy path unchanged** — same writes, same cache
hits.
- **Per-call cost:** one `getName()` / string check inside
`wrapModelCall`. Sub-microsecond.
- **Failure mode improvement:** users running `createDeepAgent` with an
Anthropic primary + non-Anthropic fallback no longer get a 400 from the
fallback provider on swap.
## Changeset
`patch` bump on `deepagents`.
## Out of scope
`request.modelSettings.cache_control` (written by LangChain core's
`anthropicPromptCachingMiddleware`) can also leak across the swap.
Mentioned in #550 for tracking — that's a LangChain-side fix, not
deepagents.
---------
Co-authored-by: Anton Nakaliuzhnyi <anakaliuzhnyi@gipartners.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> A
Anton Nak committed
18557db7bbdf92052ed5f994512fb70e11989e69
Parent: 1ca6dc9
Committed by GitHub <noreply@github.com>
on 6/2/2026, 9:04:23 PM