Anthropic has announced Claude Sonnet 5.5, the second model in the Claude 5.5 family. The easiest line to skim past is also the one that reaches your invoice most directly: the token price has not changed from Sonnet 5.
What changed
- Pricing held: $2 per million input tokens, $10 per million output tokens, $0.20 per million cache-read tokens.
- Cost per task down by up to 30%, because the model needs fewer tokens to finish the same work.
- Output generated more than 30% faster — the fastest Sonnet model so far.
- Terminal-Bench 4.0, an agentic coding evaluation: 70.6% against Sonnet 5's 10.3%.
- GDPval-AA, which tests real work across 44 occupations: 1844, exactly two points below Opus 5.5.
- The first Sonnet model to finish Pokémon Red working only from screenshots.
Why the number to read is tokens, not the score
The running cost of an AI feature is a product of two terms: tokens times price. A benchmark tells you what the model can do; the invoice depends on the first term. This time Anthropic held the price, so all of the saving sits in the model spending fewer tokens.
Early testers reported very different reductions, and that spread is the interesting part. Balyasny Asset Management, running a suite of 2,441 finance tasks, measured about 121k tokens per answer against Sonnet 5's 497k. Base44 built 118 real apps at an average of 3.6 iterations each instead of 7.7. Lovable saw a third fewer tool calls. Slack measured about 14% fewer output tokens without changing a single prompt. Box recorded 12% fewer total tokens and 2.4× the speed.

The gap between 76% and 12% is not a contradiction. Each figure belongs to a different kind of work, and the largest savings land where the older model used to over-search or repeat steps. Which makes the only figure you can plan with the one you measure on your own mix of requests.
Effort is a business parameter
Both models expose an effort level: set it low and Claude answers quickly on fewer tokens; set it high and it reasons longer and checks its work more thoroughly. The default is Medium in the Claude apps and Claude Code, High on the Claude Platform.
On Anthropic's own figures, at Low or Medium effort Sonnet 5.5 beats Sonnet 5's best score on several evaluations for about a tenth of the cost per task. Which argues for setting effort per job rather than once for the whole system: high where a mistake is expensive, low where the task repeats thousands of times a day at low risk. That is the same logic as the confidence threshold we wrote about earlier.
Tiering models, not picking one
Anthropic positions Sonnet 5.5 as a faster, cheaper complement to Opus 5.5 rather than a replacement: Opus 5.5 for complex work needing sustained judgement, Sonnet 5.5 for well-scoped tasks. One early tester put the division of labour neatly — let Opus 5.5 set the architecture and the overall framework, then hand the implementation to Sonnet 5.5.
Anyone who has designed a system will recognise the shape: put the expensive part where the decisions are, the cheap part where the volume is. It works only when the process is clear first, not when the model is good enough first — which is exactly the architectural gap that keeps many AI investments from reaching profit.
Three technical points before you switch
- If you run Sonnet with thinking off, you must move to the between_tools setting before switching to Sonnet 5.5. Anthropic has a migration guide for this.
- The model's thinking is now tied to the account that produced it. Moving a conversation between accounts — including switching accounts mid-session in Claude Code — behaves differently than before; see the preserved thinking documentation.
- Higher-risk cybersecurity requests visibly fall back to Sonnet 5. Finding and fixing bugs in your own code as part of routine development is unaffected.
The model id on the Claude Platform is claude-sonnet-5-5. It is available on Amazon Web Services, Google Cloud and Microsoft Azure, and supports zero data retention.
For companies in Vietnam
If you are already running an AI feature at volume — classifying customer requests, extracting fields from documents, answering first-line questions, labelling transactions — the first thing worth doing is not switching models but measuring your baseline: tokens per request, steps per task, and the share that has to go to a person. Without those three numbers there is no way to tell what a model change or a lower effort setting is worth.
And as always, the part that decides the value sits outside the model: confidence thresholds, stopping points for people, and a traceable log. If you are building flows like these, see our automation and AI page.
Sources: Anthropic — Introducing Claude Sonnet 5.5 · Sonnet 5.5 System Card · Migration guide
