Skip to content
Skip to main content
A single penny next to a long curling strip of blank receipt paper on green felt, a visual for Claude Haiku 5.5's cheap per-token price versus the longer token bill it runs up compared with Haiku 4.5
9 min readBy Carlos Aragon

Claude Haiku 5.5 vs Haiku 4.5: What It Really Costs

Claude Haiku 5.5 is 10x cheaper per token than Haiku 4.5, but on my real tasks it came out 75% cheaper, not 90%. The same text turns into 35-39% more tokens, adaptive thinking adds output on open-ended work, and any request over 100,000 input tokens pays 5x the rate. Set effort: "low" on simple routes and keep prompts under the threshold, and the savings go back up to 87%.

What Haiku 5.5 Actually Cost Me on Four Real Tasks

Haiku 5.5 shipped on October 7. Two days later I had a migration question on my desk, because Haiku runs most of the cheap plumbing in my automations: lead triage, field extraction, summaries that feed n8n workflows. The launch number is $0.10 in and $0.50 out per million tokens, against $1 and $5 on Haiku 4.5. Ten times cheaper on paper.

So I ran the same four tasks I actually run in production, three times each, on Haiku 4.5 and on Haiku 5.5. Median tokens and cost per call, priced at the published API rates:

TaskHaiku 4.5Haiku 5.5 (default)Haiku 5.5 (effort low)
Classify a support email506 in / 4 out · $0.00053598 in / 4 out · $0.00006598 in / 4 out · $0.00006
Extract lead fields to JSON553 in / 68 out · $0.00089669 in / 65 out · $0.00010669 in / 65 out · $0.00010
Summarize a 1,500-word article3,140 in / 206 out · $0.004174,143 in / 1,666 out · $0.001254,143 in / 226 out · $0.00053
Fix a buggy JS function513 in / 73 out · $0.00088601 in / 284 out · $0.00020601 in / 122 out · $0.00012
All four tasks$0.00647$0.00161 (−75%)$0.00081 (−87%)

Every Haiku 5.5 run passed its check (exact label, correct JSON fields, a JS fix that passes node tests, exactly five bullets). Haiku 4.5 missed the bullet count once in three summaries. So quality isn't the story here. The story is that you get roughly 4x cheaper by default and 8x cheaper with one parameter, not 10x for free.

Why Does Haiku 5.5 Use More Tokens Than Haiku 4.5?

Two separate reasons, and they hit different parts of the bill.

The tokenizer counts the same text as more tokens.I sent identical text to both models and subtracted each model's fixed request overhead. A 1,500-word article went from 2,689 tokens on Haiku 4.5 to 3,619 on Haiku 5.5, which is 34.6% more. A 360,000-character pile of my own blog posts went from 89,122 to 123,722, or 38.8% more. That lines up with Anthropic's note that the newer tokenizer produces about 30% more tokens, maybe a bit worse for my content. You pay that on every input token, cached or not.

Adaptive thinking spends output on open-ended tasks. Classification and extraction didn't trigger any thinking. The summary did: a median 1,448 thinking tokens before roughly 220 tokens of answer, 1,666 output tokens in total, against 206 on Haiku 4.5. The code fix thought for about 211 tokens. Output is the expensive side, so this is where the "90% cheaper" headline leaks most.

It costs time too. The default-effort summary took a median 6.0 seconds against 3.0 on Haiku 4.5. If you put Haiku in a voice agent or anywhere a user is waiting, that doubling matters more than the fraction of a cent. I covered the latency side of effort in which Claude effort level to use, and the same logic applies here.

What Does That Look Like on a Monthly Bill?

Per-call fractions of a cent are hard to feel, so here are my measured medians multiplied out to a realistic month for a small agency stack:

Monthly volume            Haiku 4.5   Haiku 5.5 default   Haiku 5.5 low
100,000 email triages      $52.60          $6.20              $6.20
 20,000 lead extractions   $17.86          $1.98              $1.98
 10,000 article summaries  $41.70         $12.47              $5.27
-------------------------------------------------------------------
Total                     $112.16         $20.65             $13.45

Look at where the difference comes from. The short calls hit close to the advertised 90% on their own, because there's almost no output and no thinking: the only drag is the tokenizer. The summaries are where default effort gives back money, $12.47 against $5.27 for the same five bullets. If your Haiku traffic is mostly long-ish generation, effort is the setting that decides whether the upgrade saves 70% or 87%.

The other thing this table hides is the long tail. One request a day that crosses 100K at default effort costs as much as a few hundred triage calls, so a pipeline that occasionally ingests a huge PDF can quietly dominate the bill. I'd log input tokens per request and alert on anything above 90K from day one.

What Is the Haiku 5.5 100K Pricing Threshold?

Haiku 5.5 is the one current Claude model with length-based pricing. Every other model from the 4.6 generation on bills the full 1M context at one rate. Haiku 5.5 has two price lists:

  • Prompt up to 100,000 tokens: $0.10 input, $0.50 output, $0.01 cache reads, $0.125 five-minute cache writes.
  • Prompt over 100,000 tokens: $0.50 input, $2.50 output, $0.05 cache reads, $0.625 five-minute cache writes.

Three details in Anthropic's pricing docs matter more than the prices themselves. The prompt length counts all input, including cache reads and cache writes. The higher rate applies to the whole request, output included. And each request is priced on its own, so a long agent conversation can be cheap for 30 turns and then go 5x on turn 31 when the history crosses the line.

Here's where the tokenizer and the threshold combine. My 89,122-token document was comfortably under 100K on Haiku 4.5. On Haiku 5.5 it was 123,722 tokens, so it landed in the expensive tier:

Same 360,000-character document, input cost only

Haiku 4.5   89,573 tokens  x $1.00/M  = $0.0896
Haiku 5.5  124,246 tokens  x $0.50/M  = $0.0621   <- over 100K tier
Haiku 5.5  same tokens, split into two chunks
           under 100K each x $0.10/M  = $0.0124   <- 5x cheaper

Rule of thumb: any prompt above about 72,000 tokens in your Haiku 4.5 logs will cross 100K on Haiku 5.5. That's 100,000 divided by the 1.39x I measured. If you budget from old logs, you'll underestimate exactly the requests that cost the most.

Is Haiku 5.5 Still Cheaper Than Haiku 4.5 Over 100K?

Yes, but barely enough to notice. Over the threshold Haiku 5.5 charges half of Haiku 4.5's per-token rate, and you send 39% more tokens. On my long document that worked out to 31% cheaper on input, against 86% cheaper had the same request stayed under 100K.

Output in the high tier is $2.50 per million, half of Haiku 4.5's $5. But if the task triggers thinking you send 5-8x the output tokens, so the output line ends up costing more than it did on Haiku 4.5. On a long document the input still dominates: a 124K-token summary with 1,666 output tokens comes to about $0.066, against about $0.091 on Haiku 4.5. That's 27% cheaper, a long way from the 90% you were promised.

So I wouldn't switch a long-context pipeline over blind. Measure it first. If you're working with big documents anyway, my notes on the Claude 1M context window cover when one giant prompt beats chunking. On Haiku 5.5, chunking usually wins on price.

How Do I Keep Claude Haiku 5.5 Cheap?

  1. Set effort per route.Effort low took the summary from 1,666 output tokens to 226 and from 6.0 to 1.4 seconds. That's faster than Haiku 4.5. It passed every check. One of the three low-effort summaries still thought for 967 tokens, so budget for some variance.
  2. Re-count your prompts on the new model. Run your 20 biggest real prompts through count_tokens with claude-haiku-5-5. Don't reuse Haiku 4.5 numbers for budgets, alerts or chunk sizes.
  3. Chunk under 100K and merge. Size chunks around 90K measured on Haiku 5.5 to leave room for instructions and the response history. A map-reduce over two chunks costs 5x less input than one oversized call.
  4. Cache and batch, knowing their limits. Cache reads at $0.01 per million are almost free, and the Batch API halves both tiers. Neither changes which tier a request is in, because cached tokens still count toward the 100K. My breakdowns of prompt caching savings and the Batch API still apply, with that caveat.
  5. Watch long agent histories. Compact or trim before the conversation crosses 100K, not after. The broader playbook is in AI agent cost optimization.
import anthropic

client = anthropic.Anthropic()

# 1. Know which tier you're in before you send
n = client.messages.count_tokens(
    model="claude-haiku-5-5",
    system=SYSTEM,
    messages=messages,
).input_tokens
if n > 95_000:
    raise ValueError(f"{n} tokens: chunk it, this would bill at the >100K rate")

# 2. Low effort for extraction / summaries
resp = client.messages.create(
    model="claude-haiku-5-5",
    max_tokens=2000,
    system=SYSTEM,
    messages=messages,
    output_config={"effort": "low"},
)

Should You Switch From Haiku 4.5 to Haiku 5.5?

For short, high-volume calls, switch now and set effort low. Classification and extraction were 88-89% cheaper per call with identical output, and that's most of what a Haiku-tier model does in an automation stack.

For long-document work, test with your own documents first. If your prompts sit between 72K and 100K Haiku 4.5 tokens, you're in the band that crosses the line after migration.

One tooling gotcha. Claude Code 2.1.292 still logs Haiku 5.5 as an unrecognized model, and its MAX_THINKING_TOKENS=0 setting did not stop Haiku 5.5 from thinking in my runs (1,264 thinking tokens on the summary). The Models API lists thinking: disabledas supported on Haiku 5.5, so on the raw API you can turn it off. I just didn't get to measure that path, so effort low is what I'd reach for today.

How I Tested

The API org I'd normally bench on was out of credits that morning, so I ran every call through Claude Code in headless mode (claude -p --output-format json) on my Max plan with a one-line system prompt, no tools, no MCP servers and no settings files. That leaves about 450-525 tokens of fixed overhead per request, which is included in the input numbers above and subtracted for the tokenizer comparison. Token counts come from the usage block of each response. Costs are my math at published API rates, with all input priced at the uncached base rate. Haiku 4.5 ran without thinking, matching its API default. Three runs per task per configuration, medians reported, 9 October 2026.

Haiku 5.5 Pricing FAQ

How much does Claude Haiku 5.5 cost?

For prompts up to 100,000 tokens: $0.10 per million input tokens, $0.50 per million output tokens, $0.01 per million for cache reads. For prompts over 100,000 tokens: $0.50 input and $2.50 output. The Batch API takes 50% off both tiers.

Is Claude Haiku 5.5 really 90% cheaper than Haiku 4.5?

Per token, yes. Per task, no. In my test of four tasks it was 75% cheaper at default effort and 87% cheaper at effort low, because the same text uses 35-39% more tokens and adaptive thinking adds output on open-ended tasks.

What counts toward the Haiku 5.5 100K pricing threshold?

All input tokens in the request, including cache reads and cache writes. Each request is priced on its own: if its prompt is over 100,000 tokens, the whole request pays the higher input and output rates, even if most of it was a cache hit.

Does Claude Haiku 5.5 use more tokens than Haiku 4.5?

Yes. The same English text measured 34.6% more input tokens on a 2,700-token article and 38.8% more on an 89,000-token document. Output also grows on open-ended tasks when adaptive thinking kicks in.

Should I use effort low on Claude Haiku 5.5?

For classification, extraction and summarization, yes. Effort low passed every check in my test, cut summary output by about 85% and was faster than Haiku 4.5. Keep the default for multi-step reasoning or code you can't verify automatically.

Want a Real Number Before You Migrate?

I'll run your actual prompts through both models, find the routes that cross 100K, and set effort and chunking so the bill drops the way the launch post promised. Send me what you're running.

Prices and the 100K threshold rules quoted from Anthropic's pricing page as of 9 October 2026. Model capabilities (1M context, 128K output, effort levels, adaptive and disabled thinking) read from the Models API entry for claude-haiku-5-5 the same day. Benchmarks run 9 October 2026, three runs per cell, medians shown.

Related Posts