Claude Haiku 5.5 Launches With a 90% Price Cut: Pricing, Benchmarks and the Catch

On this page
Claude Haiku 5.5 is Anthropic’s new small AI model, released on October 7, 2026, and it costs $0.10 per million input tokens and $0.50 per million output tokens for requests under 100,000 tokens. That is a 90% cut from Haiku 4.5’s list price, and it matches OpenAI’s GPT-6 Luna to the cent.
I have been watching model pricing for a few years now, and I can’t remember a week where the cheap end of the market moved this fast. Two of the biggest labs now sell their entry-level models at an identical price. That is not a coincidence. It is what a price war looks like when both sides can see each other’s rate cards.
The headline number deserves a closer look, though. A 90% cheaper token is not the same thing as a 90% cheaper bill, and the early independent numbers make that gap very clear. Here is what Anthropic announced, what the benchmarks say, and where the catches are.
What Anthropic announced with Claude Haiku 5.5
Haiku is the smallest and fastest tier in Anthropic’s lineup, sitting below Sonnet and Opus. Haiku 5.5 is the first update to that tier in nearly a year, and Anthropic describes it as its cheapest, fastest and most capable small model so far. It is available on Anthropic’s own platform under the model ID claude-haiku-5-5, and through Amazon Web Services, Google Cloud and Microsoft Azure.
Three other things shipped alongside it:
- Cheaper Sonnet caching. Cache-read pricing on Claude Sonnet 5.5 was halved, from $0.20 to $0.10 per million tokens. Anthropic estimates this makes typical agent workloads about 20% cheaper.
- API credits for subscribers. Max 5x subscribers get $100 a month in API credits, Max 20x gets $200, and Team plans get up to $500 pooled across users. The credits work on any model on the platform.
- Computer and browser control in the SDKs. The Python and TypeScript SDKs gained beta tools for operating browsers and computers.
Reporting from MarkTechPost puts the context window at 1 million tokens with up to 128,000 tokens of output, which is a lot of room for a model in the budget tier.
Claude Haiku 5.5 pricing, tier by tier
This is the part that needs a table, because Haiku 5.5 does something Haiku 4.5 never did: it charges two different rates depending on how big your request is.
| Price per million tokens | Haiku 5.5 (under 100K) | Haiku 5.5 (over 100K) | Haiku 4.5 |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 |
| Output | $0.50 | $2.50 | $5.00 |
| Cache reads | $0.01 | $0.05 | $0.10 |
| Cache writes (5 min) | $0.125 | $0.625 | $1.25 |
So the 90% figure applies only to shorter requests. Go past 100,000 tokens and the discount drops to 50%. Anthropic says roughly 90% of Haiku 4.5 requests already fell under that threshold, and once you factor in a new tokenizer that uses somewhat more tokens for the same text, the company’s own estimate is that a typical workload ends up about 75% cheaper. Batch processing takes a further 50% off.
I ran a simple example to make it concrete. Take a support bot that processes 10 million input tokens and 2 million output tokens a month, all in short requests. On Haiku 4.5 that was $10 plus $10, so $20. On Haiku 5.5 it is $1 plus $1, so $2, before the tokenizer difference nudges it up a little. For a hobby project that is pocket change either way. For a company running billions of tokens a day, it is a line item that just lost a zero.
For context, VentureBeat’s comparison of list prices has GPT-6 Luna at the same lower-tier rates as Haiku 5.5, with its higher tier starting at 272,000 input tokens instead of 100,000. Google’s Gemini 3.5 Flash-Lite sits at $0.30 input and $2.50 output. Alibaba’s Qwen3.7 Flash undercuts everyone at $0.03 and $0.13 for small inputs.
How Haiku 5.5 performs on benchmarks
Anthropic’s published numbers show a huge jump over Haiku 4.5 and a lead over GPT-6 Luna on the tests it chose to highlight. Keep in mind these are vendor-reported figures and have not been independently verified.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| OSWorld 2.1 (computer use, offline subset) | 72.4% | 15.7% | 48.9% | 83.9% |
| GDPval-AA v2.1 (knowledge work) | 1,620 | 735 | 1,437 | 1,840 |
| Terminal-Bench 4.0 (command line) | 39.2% | 0.0% | 16.4% | 70.6% |
The computer-use result is the one that caught my eye. Going from 15.7% to 72.4% on OSWorld in a single generation means the small model can now click through real software with some reliability, which is exactly the kind of grunt work you want a cheap model doing.
There is fine print here too. Haiku 5.5 is the first Haiku with adjustable effort levels, and it defaults to medium. That 39.2% Terminal-Bench score was recorded at maximum effort. At the default setting it is closer to 20%. And outside Anthropic’s own chart, Artificial Analysis lists Z.ai’s GLM-5.3-Flash at 1,647 on GDPval-AA, a notch above Haiku 5.5.
Early customers quoted by Anthropic sound happy. Box reported an 11-point improvement over Haiku 4.5 at about half the latency, and Asana said task-completion latency fell by more than 30%. Those are statements supplied through the vendor, so I would weigh them accordingly.
The catch: cheaper tokens, more of them
Think of it like a car with a very low price per litre of fuel and a large engine. What matters is the cost of the trip, not the price on the pump.
Neowin, citing testing by Artificial Analysis, reported that Haiku 5.5 averaged about 162,000 output tokens per task compared with roughly 50,000 for GPT-6 Luna. The average cost worked out at about $0.21 per task for Haiku 5.5 against $0.07 for Luna. Same sticker price, three times the bill on that particular test suite.
That does not make Haiku 5.5 a bad deal. It scored higher on those tasks, and the effort control exists precisely so developers can cap how much thinking the model does. But it does mean the honest comparison is cost per finished job, and you only find that out by running your own workload on both. I would treat any “90% cheaper” claim as a starting point for a test, not a budget forecast.
Why the cheap tier matters so much now
A year ago, small models were mostly used for simple chat and classification. Today they are the workers inside AI agent systems. A large model plans the job, then hands off dozens of smaller steps, such as summarizing a document, classifying a ticket or querying a database, to something faster and cheaper. Anthropic positions Haiku 5.5 for exactly that supporting role, and still recommends Sonnet and Opus for demanding coding.
When a single agent task can trigger hundreds of small-model calls, the price of the small model sets the price of the whole system. That is why this tier has become the battleground. It also connects to a trend we covered recently, where Anthropic measured how much of its own research work Claude now automates. Work at that scale only makes economic sense if the high-volume steps cost almost nothing.
One more detail worth noting: Haiku 5.5 ships with tighter cybersecurity safeguards than Haiku 4.5. Penetration testing is blocked by default, and organizations that need broader security or biology access have to apply through Anthropic’s verification programs. That fits the cautious mood across the industry since OpenAI said GPT-6 Astra had crossed its critical cybersecurity threshold. Capable models are getting cheaper and more locked down at the same time.
Who should switch to Haiku 5.5?
- Already on Haiku 4.5: Test it this week. If most of your requests are under 100,000 tokens, the savings are likely real, but re-check your token counts because the tokenizer changed.
- Running agents on Sonnet 5.5: You benefit even without switching, thanks to the cache-read cut. Consider moving your simplest sub-tasks down to Haiku.
- Max or Team subscribers: The new monthly API credits are a free way to try the API without pulling out a card.
- Sending very long prompts: Do the maths first. Above 100,000 tokens you pay five times the lower-tier rate.
Frequently asked questions
How much does Claude Haiku 5.5 cost?
For requests under 100,000 tokens, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. Larger requests cost $0.50 and $2.50. Batch processing is half price in both tiers.
Is Claude Haiku 5.5 better than GPT-6 Luna?
On the benchmarks Anthropic published, yes: 72.4% versus 48.9% on OSWorld 2.1 and 1,620 versus 1,437 on GDPval-AA. But independent testing reported by Neowin found Haiku 5.5 used about three times as many output tokens per task, making it costlier per job despite the identical list price.
When was Claude Haiku 5.5 released?
Anthropic released Claude Haiku 5.5 on October 7, 2026. It is available on the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure.
Does Haiku 5.5 replace Sonnet or Opus?
No. Sonnet 5.5 still scores well ahead of it, for example 70.6% against 39.2% on Terminal-Bench 4.0. Anthropic pitches Haiku 5.5 as a fast, low-cost helper for high-volume tasks, with the larger models handling harder reasoning and coding.
My read is that the real story this week is not one model. It is that two rival labs landed on the same price within days of each other, and both will have to compete on what a finished task costs. For anyone building with these tools, that is the number to start tracking.
