Claude Sonnet 5.5: Anthropic attacks the mid-market with 30% faster and 30% cheaper
🔎 Three weeks after Opus 5.5, the 5.5 family expands
Anthropic didn't waste any time. On September 28, 2026, barely three weeks after the launch of Opus 5.5, the company unveils Claude Sonnet 5.5, the second model in the Claude 5.5 family. The promise is simple: more than 30% faster generation, up to 30% lower cost per task, and performance that comes close to the company's own frontier model.
The timing is no coincidence. The price war is raging among the major labs, and Anthropic has just demonstrated with GPT-6 Sol and Luna at half price that the mid-market has become the real battleground. Sonnet 5.5 is Anthropic's answer: keep the per-token price unchanged, but make customers pay less in actual usage.
This launch is also the second since Dario Amodei publicly called for slowing the pace of improvement of the most advanced models. Anthropic makes no secret of it: according to CNBC, the company itself specifies that the model "does not push the capability frontier". We're optimizing, not revolutionizing. And that may be exactly what the market is waiting for.
The essentials
- Released September 28, 2026: Claude Sonnet 5.5 is the second model in the Claude 5.5 family, three weeks after Opus 5.5.
- 30% faster, up to 30% cheaper per task: same price per token ($2/M input, $10/M output), but fewer tokens consumed and fewer tool calls.
- Impressive benchmarks: Terminal-Bench 4.0 at 70.6% (vs. 10.3% for Sonnet 5), GDPval-AA at 1,844, virtually tied with Opus 5.5 (1,846).
- Nuanced counter-analysis: Artificial Analysis measured the highest token usage ever recorded (~193,000 output tokens per task at max effort), putting the real cost per task at $7.60, roughly 50% more than Sonnet 5.
- Five effort levels (low, medium, high, xhigh, max), native 1M context, max output 128k tokens, June 2026 cutoff.
- Available everywhere: AWS, Google Cloud, Microsoft Azure, with zero data retention available as an option.
Recommended tools
| Tool | Main use | Price (September 2026, check anthropic.com) | Best for |
|---|---|---|---|
| Claude Sonnet 5.5 | Agentic coding, automation, general-purpose use | $2/M input, $10/M output, $0.20/M cache read | Developers and teams who want Opus-level performance without the price tag |
| Claude Opus 5.5 | Frontier tasks, maximum reasoning | $4/M input, $20/M output | Complex cases where every benchmark point counts |
| Claude Code | Terminal-assisted development | Included in Claude subscriptions / billed via API | Developers who want a code agent in the CLI, medium effort by default |
| Hostinger | VPS hosting to deploy your agents and APIs | Starting at a few €/month (check hostinger.com) | Hosting your backends that call the Claude API in production |
The numbers that matter — Terminal-Bench 4.0 multiplied by almost 7
Direct answer: yes, the performance jump compared to Sonnet 5 is massive on agentic coding benchmarks, to the point of raising the question of measurement methodology.
The number that jumps out: 70.6% on Terminal-Bench 4.0, versus 10.3% for Sonnet 5. A leap of more than 60 points in a few months is rare. On CursorBench 4.0, Sonnet 5.5 reaches 55.5% versus 34.1% for Sonnet 5 — but also 57.8% for Opus 5.5, which places the mid-range model virtually on par with the frontier model.
The most revealing detail comes from GDPval-AA: 1,844 for Sonnet 5.5, 1,846 for Opus 5.5. A two-point gap. For a model billed at half the cost per token, that's the perfect marketing signal. AA-Briefcase confirms the trend with 1,811.
Other notable scores:
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | — |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% |
| FrontierCode 1.1 Max | 46.2% | — | — |
| GDPval-AA | 1,844 | — | 1,846 |
| Humanity's Last Exam (with tools) | 64.5% | — | — |
| OSWorld 2.1 | 80.1% | — | — |
| Chartography | 61.6% | 15.6% | — |
The score of 80.1% on OSWorld 2.1 deserves a closer look: it's an agentic benchmark on a real operating system, and it's exactly the type of workload that will explode with the open approach around Nvidia's OpenShell and Sentry platform. Models capable of reliably operating an OS become the raw material of production agents.
Artificial Analysis's counter-analysis — the real cost per task goes up
Direct answer: yes, the announced 30% savings hold at reference effort, but at max effort, the model consumes so many tokens that the real cost per task increases by 50% compared to Sonnet 5.
This is the most important counter-analysis of this launch, published by Artificial Analysis. Their measurements place Sonnet 5.5 at 56 on the Intelligence Index, number 2 behind Opus 5.5 (max) — an excellent position. But the token-by-token detail is eye-opening:
- ~193,000 output tokens per task at max effort, roughly 60% more than Opus 5.5 or Sonnet 5, and about 7 times more than GPT-6 Astra.
- Real cost per task of $7.60, roughly 50% more than Sonnet 5.
- At Low or Medium effort, however, the model beats Sonnet 5's best score for about one-tenth of the cost per task — there, Anthropic's promise is kept, and then some.
In other words: Sonnet 5.5 is an extraordinarily verbose model when given the means to think for a long time. It batches its tool calls more — which reduces the number of steps, as The Decoder notes — but each step generates a huge number of tokens. VentureBeat confirms that the drop in cost per task comes from the reduction in tokens and tool calls, not from a lower API price.
Second point of caution: factual knowledge. On AA-Omniscience, Sonnet 5.5 tops out at 54% versus 66% for Opus 5.5. The good news is that its hallucination rate is lower (47% versus 59%) — it knows fewer things, but lies less about what it knows. For a coding agent, that's an acceptable trade-off. For a research agent, it's a dealbreaker, and you'd be well advised to look at the best LLMs for research.
Also worth noting: Artificial Analysis identified a bug with structured outputs in pre-release, fixed before the public launch. Reassuring as to Anthropic's quality process.
The pricing stance — same price per token, a lighter bill
Direct answer: Anthropic isn't lowering its listed prices; it's lowering your bill by making the model more efficient. That's a smarter and more sustainable strategy.
Sonnet 5.5's pricing is identical to Sonnet 5's: $2/M input, $10/M output, $0.20/M cache read. As a reminder, Sonnet 5 arrived in June at this launch price, made permanent in August instead of the initially planned $3/$15. Opus 5.5, for its part, charges double: $4/$20 per million tokens.
| Model | Input (/M tokens) | Output (/M tokens) | Cache read (/M tokens) |
|---|---|---|---|
| Claude Opus 5.5 | $4 | $20 | — |
| Claude Sonnet 5.5 | $2 | $10 | $0.20 |
| Claude Sonnet 5 | $2 | $10 | $0.20 |
In the context of the price war shaking up the market since GPT-6 Sol and Luna at half price, this approach is an indirect pressure tactic: rather than joining the downward spiral of listed prices against Gemini 3.8 Flash and OpenAI's low-cost models, Anthropic is playing the per-task efficiency card. The message to customers: "Compare your monthly bills, not our rate cards."
This is also the direct continuation of the strategy initiated with Sonnet 5, whose Opus performance at a Sonnet price we analyzed. Anthropic is methodically turning its mid-range lineup into a Trojan horse: capturing the volume of developers and companies that can't justify $20/M of output, then moving them upmarket through effort levels.
On the infrastructure side, this launch rests on considerable compute capacity — the recent deal with SpaceX for Colossus 1 and its 220,000 GPUs gives a sense of the scale of resources committed to sustaining this release cadence.
Five effort levels — the hidden lever of cost
Direct answer: the real economic innovation of Sonnet 5.5 isn't the price, it's the granular control over reasoning effort, which lets you divide your bill by ten without sacrificing quality.
Sonnet 5.5 offers five effort levels: low, medium, high, xhigh, max. The default is high on the API and medium in Claude Code. This is the most under-communicated point of the launch, and yet the most important for your wallet.
Concretely:
- At Low or Medium effort, Sonnet 5.5 surpasses the best score ever achieved by Sonnet 5 for roughly one tenth of the cost per task. This is the absolute sweet spot for common tasks: code review, simple refactoring, test generation, support agents.
- At High effort (API default), you get the announced benchmark scores, with token consumption that is already high.
- At Max effort, you reach Opus 5.5's level on many tasks, but with ~193,000 output tokens per task and $7.60 per task according to Artificial Analysis. At that level, the question to ask is simple: why not just use Opus 5.5 directly at $4/$20?
The practical recommendation is therefore clear: set the effort to Medium by default in your applications, and only go up to xhigh/max for explicitly difficult tasks. That's the difference between a controlled API bill and a nasty surprise at the end of the month.
Another specification to know: 1M token native context, max output of 128k tokens (300k on Message Batches in beta), June 2026 cutoff, and the ID claude-sonnet-5-5. As for generation speed, it is more than 30% faster than Sonnet 5 — a gain you feel immediately in Claude Code and in agentic loops where latency accumulates at every step.
What This Changes for Developers in Production
Direct answer: for 90% of coding use cases, Sonnet 5.5 makes Opus 5.5 unnecessary. The math is brutal and in favor of switching to the mid-range lineup.
If you develop with Claude on a daily basis, here's the concrete impact:
In Claude Code, the default Medium effort setting means you benefit from the speed and cost gains without changing anything in your setup. For everyday coding tasks — understanding a codebase, writing a function, fixing a bug — Sonnet 5.5 at Medium is very likely the best quality/price ratio on the market. Our comparison of the best LLMs for coding will be updated accordingly.
In your CI/CD pipelines and automated agents, the reduced number of tool calls (thanks to batching) decreases the number of steps per task. Fewer steps means less latency, fewer points of failure, and fewer tokens. This is particularly relevant for the agentic architectures described in our guide on the best LLMs for AI agents.
On the operational side, the model is deployed on AWS, Google Cloud, and Microsoft Azure, with zero data retention available. For companies with compliance constraints, that's a strong argument — and the cybersecurity safeguards are, for the first time on a Sonnet, comparable to those of frontier models.
For frontier use cases, on the other hand, Opus 5.5 still has value: GDPval-AA is indeed tied (1,846 vs. 1,844), but the gap widens on factual knowledge (66% vs. 54% on AA-Omniscience) and on tasks where the token budget isn't an issue. If you're hesitating between the two ecosystems, our Claude vs ChatGPT comparison and the monthly comparison of the best LLMs will help you decide.
And if you'd rather keep control over your data and costs, the best LLMs to run locally remain a credible alternative for sensitive tasks, along with our guide to installing a local LLM via Ollama or LM Studio.
Security and system card — less misalignment, but more opaque reasoning
Direct answer: the system card is fairly reassuring on the risk front, with one notable caveat about the legibility of its reasoning.
The RSP (Responsible Scaling Policy) evaluations conclude that Sonnet 5.5 is overall less capable than Opus 5.5 and crosses no new threshold. Misalignment risks are assessed as low. An interesting point: it's the model with the lowest propensity of all those tested to probe the boundaries of its containers — good news for anyone running agents in sandboxed environments, at a time when initiatives like Nvidia's OpenShell platform are trying to put guardrails around agent safety.
The downside comes on the interpretability side: Sonnet 5.5's reasoning is harder to read than that of previous models. For a company that invests in auditing its agents — as Anthropic itself demonstrated with its automated alignment researcher that outperforms humans — this is a tension worth watching. Performance gets optimized, transparency gets degraded.
❌ Common Mistakes
Mistake 1: Believing that "30% cheaper" applies to all use cases
What goes wrong: the 30% savings are measured at reference effort. At Max effort, Artificial Analysis measures a cost per task of $7.60, roughly 50% more than Sonnet 5, due to ~193,000 output tokens per task.
The fix: lock Medium effort as the default in your applications, and reserve xhigh/max for explicitly difficult tasks. Measure your actual cost per task, not the price per token.
Mistake 2: Removing Opus 5.5 from your stack too quickly
What goes wrong: the tie on GDPval-AA (1,844 vs 1,846) tempts you to switch 100% to Sonnet 5.5. But the 12-point gap on AA-Omniscience (54% vs 66%) shows the frontier model retains the advantage on factual knowledge.
The fix: keep task-based routing — Sonnet 5.5 for code and agentic work, Opus 5.5 for factual analysis and high-stakes decisions.
Mistake 3: Comparing price lists instead of invoices
What goes wrong: with a per-token price identical to Sonnet 5, you might conclude there's no economic gain.
The fix: compare the cost per completed task on your own workloads. That's the metric that matters, and it's exactly where Anthropic wants to help you win.
Mistake 4: Ignoring the degradation in reasoning readability
What goes wrong: the system card notes that the reasoning is less legible than that of previous models, which complicates auditing your agents.
The fix: if auditability is critical for you, set up structured traces upstream and downstream of model calls, and test the structured outputs (the pre-release bug has been fixed, but validate on your own cases).
❓ Frequently Asked Questions
What is the exact price of Claude Sonnet 5.5?
$2 per million input tokens, $10/M output, and $0.20/M cache read (September 2026, check on anthropic.com). This is identical to Sonnet 5 and half the price of Opus 5.5 ($4/$20). The announced 30% savings comes from reduced consumption, not the listed price.
Does Sonnet 5.5 replace Opus 5.5?
Not entirely. On agentic coding and GDPval-AA, the two are virtually tied, and Sonnet 5.5 costs half as much. But Opus 5.5 retains a clear advantage on factual knowledge (66% vs 54% on AA-Omniscience) and on Max effort tasks. Task-based routing remains the best strategy.
What is the difference from Sonnet 5?
Over 30% faster generation speed, up to 30% lower cost per task, and massive gains on benchmarks: Terminal-Bench 4.0 jumps from 10.3% to 70.6%, CursorBench 4.0 from 34.1% to 55.5%. Add to that five effort levels and frontier-level cybersecurity safeguards, a first for a Sonnet.
Is the model available on the clouds?
Yes, Sonnet 5.5 is deployed on AWS, Google Cloud, and Microsoft Azure from launch, with zero data retention available as an option. The model ID is claude-sonnet-5-5, with native 1M token context and 128k max output tokens.
Should we worry about safety risks?
The system card judges misalignment risks to be low, and the model does not cross any new RSP threshold. It is also the model least inclined to probe the limits of its containers. The main point of vigilance is the growing opacity of its reasoning, which complicates auditing agents in production.
Is Claude Haiku 5.5 coming soon?
Anthropic has announced Claude Haiku 5.5 for the coming weeks. Logically, it will inherit the efficiency gains of the 5.5 family at an even lower price point, which could reshuffle the deck in the budget model segment against GPT-6 Sol and Gemini 3.8 Flash.
✅ Conclusion
Claude Sonnet 5.5 is not a frontier model, and Anthropic says so itself — it's better than that: it's the most cost-effective model in its lineup, provided you master the effort slider. At Medium, it buries Sonnet 5 for a tenth of the cost; at Max, it catches up with Opus 5.5 but sends the token counters soaring. Try it today with the default Medium effort, and check out our comparison of the best LLMs for coding to see how it stacks up against the competition.