Skip to content
Skip to main content
Two identical stacks of cream index cards with one orange-edged card pulled out of the right stack and a brass magnifying glass beside it — a Claude prompt cache miss traced to the one part of the request that changed
8 min readBy Carlos Aragon

Claude Prompt Cache Miss? What Cache Diagnostics Shows

When your Claude prompt cache stops hitting, add a diagnosticsobject to every request and pass the previous response's id. The API then tells you exactly what changed: the model, the system prompt, the tools or the message history. It went GA on 23 September, no beta header needed. This morning I broke a 21,000-token cache five different ways on purpose. Three of them threw away every cached token, one of them was a single trailing space, and diagnostics named the cause every time.

Why Does a Claude Prompt Cache Miss Silently?

Prompt caching only works when the start of your request is byte-for-byte identical to a recent one. Change anything before the breakpoint and the API quietly bills the whole prefix again, at the cache write rate. Until last week the only symptom was usage.cache_read_input_tokens dropping to zero. No error, no hint.

That's expensive in a way people underestimate. A cache read costs 0.1x the base input price on most models and a 5-minute write costs 1.25x, so a miss on a big prefix costs about 12.5 timeswhat the hit would have. If you want the full break-even math, it's in my Anthropic prompt caching cost breakdown. On Opus 5.5 and the Fable 5.1 models the read discount is even deeper, which makes every miss hurt more.

A silent cache miss isn't a failure you'll see in logs. It's a line on the invoice.

How Do You Turn On Cache Diagnostics?

One field. On the first turn you opt in with a null id. On every turn after that you pass the id of the previous response. Here's the Python loop I use:

prev_id = None
for user_msg in turns:
    messages.append({"role": "user", "content": user_msg})
    r = client.beta.messages.create(
        model="claude-sonnet-5",
        max_tokens=1024,
        cache_control={"type": "ephemeral"},
        system=SYSTEM,
        tools=TOOLS,
        messages=messages,
        diagnostics={"previous_message_id": prev_id},
    )
    d = r.diagnostics
    if d and d.cache_miss_reason:
        log.warning("cache miss: %s (%s tokens)",
                    d.cache_miss_reason.type,
                    getattr(d.cache_miss_reason, "cache_missed_input_tokens", None))
    messages.append({"role": "assistant", "content": r.content})
    prev_id = r.id

Three details from the docs that matter in practice:

  • Send the object on every turn. The API only stores a fingerprint for requests that include diagnostics. Skip it once and the next turn gets previous_message_not_found.
  • Claude API only.It's not on Amazon Bedrock, Google Cloud, Microsoft Foundry or Claude Platform on AWS yet.
  • No prompt text is stored.The fingerprint is hashes and token estimates, scoped to your workspace, and it expires quickly. It's ZDR eligible.

When streaming, diagnostics arrives on the message_start event, so you can log it before the first token.

What Happened When I Broke the Cache on Purpose?

I built a support-agent prompt on Claude Sonnet 5: a system prompt with 400 refund policies (about 21,000 tokens), two tools, and top-level automatic caching. Turn 1 wrote the cache. Then I fired turns that each changed one thing and compared against the same stable turn 2.

What I changedCache read / writecache_miss_reason
Nothing, just appended a turn20,969 / 55null
Current time at the top of the system prompt0 / 21,090system_changed
Same two tools, reversed order0 / 21,079tools_changed
One trailing space on the first user message0 / 21,080messages_changed
tool_choice set to none21,024 / 55null
Made-up previous_message_id21,079 / 0previous_message_not_found

The timestamp one is the classic. Somebody adds “Today is {date}” to the system prompt so the agent knows the date, and caching dies for the whole app. The reversed tools are the sneaky one: I've seen tool lists built from a dict or a Set of enabled integrations, where the order depends on which ones loaded first.

The trailing space is the one that should scare you. That's what happens when your app trims or re-renders stored history before sending it back, or when an assistant turn is saved as plain text and resent without its original blocks. One invisible character threw away 21,000 cached tokens, and diagnostics pointed straight at the message history.

Two notes on reading the output. cache_missed_input_tokenscame back as 21,675 on all three misses, while the real write was around 21,080. The docs say it's estimated from byte lengths, so use it for size, not for billing. And the tool_choicechange didn't cost me a miss on this run. The docs list tool_choice among the parameters that return unavailablewhen they change, so don't read my one test as a promise.

How Do You Read Diagnostics Alongside Usage?

Diagnostics answers “did my request change?” Usage answers “did the cache hit?” You need both, and the combination tells you where to look:

  • Null diagnostics, high cache reads: working. Leave it alone.
  • Null diagnostics, zero cache reads: your request was stable but the entry expired. The default TTL is 5 minutes. For agents that wait on humans, use the 1-hour TTL.
  • A *_changed type, zero cache reads: your bug. Fix what the type names.
  • {"cache_miss_reason": null}: the comparison hadn't finished when the response went out. Inconclusive; check the next turn.
  • unavailable: no comparison. Usually a changed thinking, output_config, beta header set or similar parameter, or a divergence too deep in a very long conversation.

Diagnostics reports the earliestdivergence only. If your system prompt and your tools both change, you'll see system_changed, fix it, and only then find out about the tools. Budget for two or three rounds on a messy codebase.

How Do You Fix Each Cache Miss Reason?

  • model_changed: a router, fallback or A/B test switched models mid-conversation. The cache is per model. Pin the model for the life of a conversation and only route at the start.
  • system_changed: make the system prompt a constant. Move the date, user name and request IDs into the first user message, after the breakpoint.
  • tools_changed: sort the tools by name and serialize schemas with sorted keys. If you genuinely need to add a tool mid-conversation, the new inline tool definitions let you do that without touching tools. I covered that pattern in changing Claude's tools mid-conversation.
  • messages_changed: treat history as append-only and send assistant content back exactly as the API returned it. If you need to shrink history, do it on purpose with the compaction API instead of trimming old turns every request.

One more thing is worth doing regardless of the reason: put an explicit breakpoint at the end of your system prompt. With only automatic caching, my one-space history edit cost the full 21,000 tokens, because the only cache entry sat at the end of the conversation. I re-ran it with cache_control on the system block. Same edit, same messages_changed reason, but the API read 20,884 tokens from cache and rewrote just 135.

A breakpoint after your static prefix turns a full-price miss into a rounding error. It's also why thinking-model migrations push you toward append-only history. On Opus 5.5, editing an earlier message can also break preserved thinking, which I wrote about in the Opus 5.5 migration notes.

Should You Leave Diagnostics On in Production?

I would. Anthropic documents it as best-effort: it never blocks or fails a request, and if the comparison isn't ready it just returns an inconclusive result. It stores no prompt text. The real cost is one extra field in your request builder and one log line.

What I do with it: log cache_miss_reason.type, cache_read_input_tokens and cache_creation_input_tokens per request, then alert when the share of *_changed misses climbs after a deploy. That catches the teammate who adds a timestamp to the prompt on a Friday before it shows up on the monthly bill. The official cache diagnostics docs have the same loop in every SDK, and the prompt caching guide covers minimum lengths and TTLs if diagnostics comes back clean and you're still missing.

Is Your Claude Bill Higher Than It Should Be?

I build and tune Claude API agents in production, and a cache audit is usually the fastest money I find. Send me what you're running and I'll tell you where the tokens are going.

Token counts and cache_miss_reason values measured 29 September 2026 against the Claude API with claude-sonnet-5, a ~21,000-token system prompt, two tools and 60-token responses; eleven requests in total. Feature behavior, response format and platform availability checked against the Claude Platform cache diagnostics docs and release notes (GA 23 September 2026) the same day.

Related Posts