Anthropic completes the Claude 5.5 family with Haiku 5.5: three models in one month and 1 million tokens of context
🔎 The newest addition that says a lot about the industry
On October 8, 2026, Anthropic launched Claude Haiku 5.5, the third model in the 5.5 family in a little over two weeks. After Opus 5.5 (September 22) and Sonnet 5.5 (September 28), the lab is completing its lineup with its smallest, fastest, and cheapest model. The roundup of the day's announcements relayed by Aidapted confirms a pace that no one sees slowing down anymore.
The headline figure: $0.10 per million input tokens, 90% less than Haiku 4.5 under 100k tokens. But the price isn't the real story.
In 16 days, Anthropic has renewed its entire frontier lineup, with a context window raised to 1 million tokens across the board. A frontier model now has a commercial lifespan of a few weeks. And the labs' strategy increasingly resembles a portfolio of models — one per task, one per budget — rather than a single "do-it-all" model.
The essentials
- Haiku 5.5 launched on October 8, 2026: $0.10/M input and $0.50/M output up to 100k tokens, i.e., 90% less than Haiku 4.5 (roughly 75% cheaper on average).
- Three 5.5 models in 16 days: Opus 5.5 (September 22), Sonnet 5.5 (September 28), Haiku 5.5 (October 8), with 1M tokens of context for the family and 128K max output for the two older siblings.
- A deliberate positioning: summaries, context compaction, classification, and a subagent role for Opus 5.5 and Sonnet 5.5.
- Day-one availability on the Anthropic API, AWS, Google Cloud, and Azure, with Python/TS SDKs featuring computer use and browser use in beta.
- Our take: at this frontier release pace, any model can be surpassed within weeks. The right question is no longer "which is the best model," but "which model for which task, at what price."
Recommended tools
| Tool | Main use case | Price (October 2026) | Best for |
|---|---|---|---|
| Claude Haiku 5.5 | High-volume tasks, subagents | $0.10/$0.50 per M tokens | Summaries, classification, compaction |
| Amazon Bedrock | Enterprise access to Claude 5.5 | Pay-as-you-go | Compliant deployments, VPC |
| APIpulse | Tracking API pricing | Free | Compare prices before committing |
| Hostinger | Hosting the apps that call the API | From a few dollars/month (check hostinger.com) | Deploy an API backend with no friction |
Haiku 5.5: what exactly is it?
It's Anthropic's smallest, fastest, and cheapest model, identified as claude-haiku-5-5 in the API, and designed for high-volume tasks — where cost per token matters more than cutting-edge reasoning.
Anthropic positions it bluntly: summarization, compaction, classification. In other words, the thankless but massive work that keeps AI pipelines running. Like its two older siblings, it inherits the 1 million token window that is becoming the family standard, and its tiered pricing — one price up to 100k tokens of context, another beyond — cements this long-document orientation.
Two details deserve your attention. First, the Python and TypeScript SDKs ship with computer use and browser use in beta: Haiku 5.5 was designed from the start to drive low-cost agents. Second, it's available immediately on AWS, Google Cloud, and Azure — the era of three-month exclusives seems to be over.
A subagent above all
Anthropic openly embraces Haiku 5.5's role as the right-hand model to Opus 5.5 and Sonnet 5.5: the expensive model reasons, the cheap model executes. This is exactly the architecture we break down in our feature on the best LLMs for AI agents.
Proof that this pattern has gone mainstream: Cognition uses Haiku 5.5 as a sidekick in Devin Fusion, which posts a FrontierCode score of 66.2 according to the figures cited by Anthropic. A $0.10 model taking part in a top-tier coding agent chain — that's the signal to read in this announcement.
Pricing: -90% under 100k tokens, and a tier that changes the math
Haiku 5.5 costs $0.10/M for input and $0.50/M for output for requests up to 100k tokens. Beyond that, expect $0.50/M for input and $2.50/M for output — five times more. Cache reads drop to $0.01/M ($0.05 beyond 100k).
| Model | Input /M tokens | Output /M tokens | Window |
|---|---|---|---|
| Claude Haiku 5.5 | $0.10 (≤100k) / $0.50 | $0.50 (≤100k) / $2.50 | 1M |
| Claude Haiku 4.5 | $1 | $5 | 200K |
| Claude Sonnet 5.5 | $2 | $10 | 1M |
| Claude Opus 5.5 | $5* | $25* | 1M |
*Opus row reference (Opus 5) according to the pricing grid compiled by APIpulse, October 2026. Check the prices in effect on Anthropic's website before making any commitment.
Under 100k tokens, Haiku 5.5 cuts Haiku 4.5's bill by a factor of ten, for both input and output. Beyond that, the discount drops to 50%: Anthropic is charging for long context, a classic business logic when the window climbs to 1M. On average, the lab announces roughly 75% savings depending on the usage mix.
Another quiet but telling move: Sonnet 5.5's cache reads have been cut in half, which works out to roughly 20% in savings on agentic workloads. The message is clear — Anthropic is optimizing its pricing for agents in production, not for demos.
API credits for Max and Team subscribers
Anthropic is also introducing a monthly API credit for subscribers: $100 for Max 5x, $200 for Max 20x, $500 for Team. Subscription and API are starting to converge.
This is no accident. The lab has already switched Fable 5 to credit-based billing, at $10 and $50 per million tokens. The underlying trend is easy to read: unbundle the pricing of frontier models, and defend margins on small models at volume. Haiku 5.5 at $0.10 is the visible face of that.
Three models in 16 days: the cadence that makes a model obsolete
A frontier model now stays "the best" for a few weeks, not a few quarters. Opus 5.5 arrived on Amazon Bedrock on September 22, 2026, Sonnet 5.5 on September 28, and Haiku 5.5 on October 8, according to the model cards on Amazon Bedrock.
| Date | Model | Highlight |
|---|---|---|
| Sep 22, 2026 | Claude Opus 5.5 | 1M context, 128K output, permanent adaptive thinking |
| Sep 28, 2026 | Claude Sonnet 5.5 | 1M context, adjustable effort, June 2026 cutoff |
| Oct 8, 2026 | Claude Haiku 5.5 | $0.10/M input, built for volume and subagents |
This is not an isolated case. GPT-5.5 and Gemini 3.1 Pro are battling for the top of the general-purpose leaderboard, DeepSeek V4 Pro and Kimi K2.6 are driving prices down from China, and Mistral released Large 4 and its 1000 billion open-weight parameters — the European bet on the frontier. Every lab now ships on a monthly cadence, sometimes biweekly.
The consequence for benchmarks is brutal: the ranking you read this morning will be outdated in three weeks. That's precisely why we maintain a monthly comparison of the best LLMs rather than an annual ranking. And that's why "what is the best model on the market" has become the wrong question.
What this means in practice
Reevaluate your stack every quarter, not every year. The model that justified $1/M at the start of the year can be replaced by an equivalent at $0.10/M in October — that's literally what just happened between Haiku 4.5 and Haiku 5.5. Those who signed annual commitments on the old pricing grid are now paying ten times the market price for the same work.
The real strategy: a portfolio of models, not one model for everything
Anthropic no longer sells a flagship model; it sells a portfolio: Opus 5.5 for hard reasoning, Sonnet 5.5 for everyday work, Haiku 5.5 for volume. Every task has its price, and every price has its model.
Adaptive thinking, moreover, blurs the line between sizes. On Sonnet 5.5, reasoning effort is adjustable from low to max; on Opus 5.5, it's always active, defaulting to medium. You're no longer buying a brain — you're buying an effort dial. Model size becomes an infrastructure parameter, like a cloud instance size.
This fleet logic is already running at scale. The Glasswing operation, where Claude Mythos discovered more than 10,000 critical vulnerabilities in a single month, illustrates the portfolio economy: masses of cheap calls coordinated by an expensive model, with a bottleneck that shifted the problem from detection to patching. That's exactly the architecture Haiku 5.5 wants to democratize — and the subagent architecture mentioned above makes it accessible to any team.
The parallel with cloud computing is striking. Nobody asks anymore "what's the best server instance": they ask which instance for which workload. AI models have just reached this stage of pricing maturity, and Anthropic is the first lab to embrace it this explicitly in its product messaging.
Haiku 5.5 in production: where it shines, where it doesn't belong
It shines wherever volume dominates: document summarization, context compaction between agent turns, ticket classification, structured extraction. Steer clear as soon as a task demands high-level multi-step reasoning — that's Opus 5.5's job, and it bills roughly 50 times more for input.
The math is unforgiving. Processing 100 million tokens of summaries costs about $10 in input with Haiku 5.5, versus $100 with Haiku 4.5 and $200 with Sonnet 5.5. On a pipeline running 24/7, the gap adds up to thousands of euros per year.
On the agent side, the profile is that of an executor: computer use and browser use in beta in the SDKs, floor pricing for short loops, and a sidekick role validated by Cognition in Devin Fusion (FrontierCode 66.2). For heavy coding — architecture, complex debugging, refactors — stick with Sonnet or Opus 5.5: our selection of the best LLMs for coding details the thresholds where the small model falls short.
Limitations to know about
First, the 100k token tier: beyond that, each token costs five times more ($0.50/M in input, $2.50/M in output). Split up or summarize your long documents before sending them.
Next, the cutoff: Sonnet 5.5 stops at June 2026, and it's safe to assume a comparable horizon for Haiku 5.5. For anything involving recent events, plan for RAG or a research-oriented LLM with web access.
Finally, an often-forgotten obvious point: when your AI pipeline scales up, your backend has to keep up. A VPS from Hostinger is more than enough to orchestrate API calls — its cost is negligible compared to the token bill.
Facing the competition: the table that matters
Under 100k tokens, Haiku 5.5 matches Gemini 2.5 Flash-Lite on input ($0.10/M) and halves GPT-5.4 nano's entry price ($0.20/M). On output, its $0.50/M still trails Flash-Lite ($0.40/M) but crushes nano ($1.25/M).
| Model | Input /M | Output /M | Positioning |
|---|---|---|---|
| Claude Haiku 5.5 | $0.10 | $0.50 | Volume + long context (1M family) |
| GPT-5.4 nano | $0.20 | $1.25 | Volume |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | Volume, previous generation |
| Claude Sonnet 5.5 | $2 | $10 | Generalist, 1M context |
| Gemini 3.1 Pro | $2 | $12 | Generalist |
| GPT-5.4 | $2.50 | $15 | Generalist |
Source: APIpulse, October 2026.
The honest reading: Flash-Lite remains about 20% cheaper on output, but it's a previous-generation model, without Anthropic's agentic ecosystem — computer use, browser use, native subagent role. Against GPT-5.4 nano, Haiku 5.5 is cheaper on both axes: the generational gap reads directly in the prices.
For everything else, the head-to-head duel remains relevant: our Claude vs ChatGPT comparison draws the dividing line between the two ecosystems, beyond per-token prices.
How to get started with Haiku 5.5 today
The model is available everywhere, immediately: Anthropic API (ID claude-haiku-5-5), Amazon Bedrock, Google Cloud, and Azure. No waitlist, no gradual rollout.
Your first call in a single command:
curl https://api.anthropic.com/v1/messages \
-H "x-api-key: $ANTHROPIC_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model":"claude-haiku-5-5","max_tokens":512,"messages":[{"role":"user","content":"Summarize the following text in three points: <your text>"}]}'
Three tips to get off to a good start. Route by task from day one: Haiku for volume, Sonnet as your default, Opus on demand — that's the portfolio logic. Enable prompt caching: at $0.01/M for cache reads on Haiku 5.5, and with Sonnet 5.5 cache reads cut in half, the optimization shows up immediately on your bill. Finally, if you're a Max or Team subscriber, a monthly API credit is now included — $100, $200, or $500 depending on your plan.
❌ Common Mistakes
Mistake 1: Paying Sonnet or Opus for high-volume work
Sending 50 million monthly classification tokens to Sonnet 5.5 means paying $2/M where $0.10/M would suffice: twenty times more expensive for results that are often identical. Solution: audit your logs, isolate the mechanical tasks, and route them to Haiku 5.5.
Mistake 2: Ignoring the 100k token tier
The appeal of 1M context hides a pricing trap: beyond 100k tokens, everything costs five times more. Solution: split, summarize, or compact before sending — which is, incidentally, exactly the kind of work Haiku 5.5 is built for.
Mistake 3: Building your entire stack on a single model
A sobering reminder: Trump ordered the blocking of Claude Fable 5 and Mythos 5, and Anthropic disabled these models for all users. A model can disappear overnight, for reasons that have nothing to do with technology. Solution: abstract away the provider, keep a tested plan B, and avoid over-engineered prompts tied to a single model.
Mistake 4: Neglecting cache reads
Haiku 5.5 cache reads cost $0.01/M under 100k tokens — ten times less than standard input. If you don't structure your prompts for caching (stable instructions at the top, variable content at the end), you're leaving on the table the 20% agentic savings that Anthropic highlighted for Sonnet 5.5.
❓ Frequently Asked Questions
Is Haiku 5.5 free?
No, it's a paid API model: $0.10/M for input, $0.50/M for output under 100k tokens. However, Max and Team subscribers receive a monthly API credit of $100 to $500 depending on the plan. To test without a credit card, check out our selection of the best free LLMs.
Can Haiku 5.5 code?
Yes, for routine tasks: Cognition uses it as a sidekick in Devin Fusion, with a FrontierCode score of 66.2. But for software architecture, complex debugging, or overhauls, stick with Sonnet or Opus 5.5 — considerably more expensive, but built for that level of reasoning.
How does it compare to GPT-5.4 nano and Gemini 2.5 Flash-Lite?
Price first: $0.10/M versus $0.20/M for input for nano, and $0.50/M versus $1.25/M for output. Then the ecosystem: computer use, browser use, and native subagent integration. Flash-Lite remains slightly cheaper on output, but it belongs to a previous generation.
Is there a local alternative?
Yes, if you can accept a quality gap and have a decent GPU. Open-weight models run on Ollama or LM Studio at a marginal cost per token. Follow our guide to installing a local LLM and our selection of the best local LLMs to choose based on your hardware.
Which Claude model should you choose in October 2026?
Three answers: Opus 5.5 if the task justifies maximum reasoning, Sonnet 5.5 as the default for everyday use, Haiku 5.5 whenever volume dominates the bill. It's the portfolio logic that Anthropic now fully embraces — and one your pipelines would do well to copy.
✅ Conclusion
By wrapping up its 5.5 family in sixteen days — Opus, Sonnet, Haiku, 1M context — Anthropic buries the question of the "best model" in favor of a far more useful one: which model for which task, at what price. The answer changes every month; our comparison of the best LLMs tracks it for you.