Sonnet 5.5 Costs the Same per Token as Sonnet 5, and Up to 30% Less per Task

Sonnet 5.5 Costs the Same per Token as Sonnet 5, and Up to 30% Less per Task
On 28 September 2026 Anthropic released Claude Sonnet 5.5, the second model in the 5.5 generation after Opus 5.5, which shipped on 22 September. The price did not move from Sonnet 5: $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens read from the prompt cache. A token is the unit of text AI models are billed by.
Even so, Anthropic says the new model costs up to 30% less per task in its testing, and writes its output more than 30% faster. The reason is not a discount. The model simply needs fewer tokens to do the same work. That is the point worth any business's attention if it pays an AI bill: the price per token does not tell you what a task costs.
Why the bill falls while the token price stays put
A model bill is the number of tokens multiplied by their price. And the token count is not set by the length of your question alone. It depends on what the model does to reach an answer: how many steps it takes, how often it calls a tool or searches, and how much it thinks before it writes. Anthropic's documentation states that thinking tokens are billed as output tokens, the more expensive kind.
Alongside the announcement, Anthropic published results from companies that tested the model before launch:
- Balyasny Asset Management: on 2,441 finance tasks covering Q&A, extraction, analysis and forecasting, Sonnet 5.5 used about 121,000 tokens per answer against 497,000 for Sonnet 5, and scored higher.
- Slack: it beat Sonnet 5 on almost all of Slack's offline Slackbot evaluations, in fewer steps and with about 14% fewer output tokens, without any change to the prompts.
- Lovable: in its coding evaluations it needed a third fewer tool calls and roughly half the shell runs to finish a task.
- Box: it was more accurate than the previous model, 2.4 times faster, and used 12% fewer total tokens.
Notice the range: from 12% fewer tokens to roughly three quarters fewer, depending on the kind of work. These are the companies' own tests as published by Anthropic, not independent measurements. In practice, the real saving in your business can only be known by trying it on your own tasks.
Sonnet 5.5 or Opus 5.5: the difference in price and capability
Opus 5.5 charges twice as much for input and output, while cache reads cost the same, and on most published benchmarks the gap is small. But Anthropic itself gives each model a different job.
| Item | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| API model ID | claude-sonnet-5-5 | claude-opus-5-5 |
| Price (input / output per million tokens) | $2 / $10 | $4 / $20 |
| Cache write (5 minutes) | $2.50 | $5 |
| Cache read | $0.20 | $0.20 |
| Batch API (input / output) | $1 / $5 | $2 / $10 |
| Context window / max output | 1M tokens / 128K tokens | 1M tokens / 128K tokens |
| Default effort on the API | high | medium |
| Terminal-Bench 4.0 — terminal coding | 70.6% | 66.4% |
| CursorBench 4.0 — tasks from real coding sessions | 55.5% | 57.8% |
| OSWorld 2.1 — computer use | 80.1% | 81.8% |
| GDPval-AA — professional work across 44 occupations (rating) | 1844 | 1846 |
| Humanity's Last Exam — with tools | 64.5% | 67.7% |
Figures come from Anthropic's announcement and its pricing and models documentation. The Opus 5.5 Terminal-Bench score is its highest, at Xhigh effort. For comparison, Sonnet 5 rated about 1449 on GDPval-AA.
The numbers are close, but the same announcement says benchmarks capture only one facet of capability, and that in Anthropic's own testing and that of external testers Opus 5.5 remains clearly stronger at complex, open-ended work that needs sustained judgment. Sonnet 5.5 is at its best on well-scoped everyday tasks, bug fixing, and producing documents, slides and spreadsheets.
Our reading, in operational terms: answering a customer query under clear rules, extracting invoice data, summarising a meeting, or drafting a deck from a fixed template are Sonnet jobs. Designing a new system, an open-ended analysis that weighs options, or a decision that rests on long and tangled context are where Opus may earn its price. We compared Opus 5.5 with the top-tier Fable 5.1 in detail in an earlier article.
Effort: the setting that moves the bill more than the model name
Sonnet 5.5 runs at five effort levels: low, medium, high, xhigh and max. The higher the level, the longer the model works and the more it checks itself, so the cost per task rises and the score usually does too. The default is medium inside the Claude apps and Claude Code, and high when you call it through the API.
Anthropic says that on several benchmarks Sonnet 5.5 at low or medium effort beat Sonnet 5's best score for about a tenth of the cost per task. On FrontierCode, which measures whether a code change could be merged without human edits, it scored 10 points above Sonnet 5 at the same high setting, for about one fifteenth of the cost per task.
One note in the announcement deserves a pause: on FrontierCode, Sonnet 5.5 scored lower at max effort than at xhigh, because it more often ran Claude Code's code-review skill, which splits the review across many subagents; in two cases examined by Cognition, this led to a timeout or to edits beyond the task. The highest effort is not always the best result, but it usually costs the most.
The migration guide also warns that the levels were recalibrated, so a given level does not produce the same amount of thinking as it did on Sonnet 5. It recommends starting at high for general work, at medium for well-specified coding and multi-step tool tasks, and at medium or low for chat and anything latency-sensitive. In other words, a system that calls the model through the API without setting a level will run at high, and may pay more than its routine tasks need.
Before your developer changes the model name
Switching models looks like a one-line change, but the migration guide lists changes that can stop a working integration or alter its bill:
- Turning thinking off has changed: systems that disable thinking on Sonnet 5 with
thinking: {"type": "disabled"}will get a 400 error from Sonnet 5.5. The replacement is the newbetween_toolssetting, which is accepted only at low, medium and high effort. - Thinking is on by default: a request that sets no thinking field runs with adaptive thinking on Sonnet 5.5. Anyone moving from Sonnet 4.6 or earlier, or from Haiku 4.5, where thinking was off by default, will now pay for thinking tokens on requests that never set a thinking field.
- The same text, more tokens, if you come from generation 4: Sonnet 5.5 uses Sonnet 5's tokenizer, and the documentation says the same text produces about 30% more tokens than on Sonnet 4.6, Sonnet 4.5 and Haiku 4.5, depending on the content. The per-task saving Anthropic announces is measured against Sonnet 5, not against generation 4.
- Response shape: a response can begin with a thinking block before the text, so code that reads the first element of the reply as the text will break.
- Rejected settings: fixed thinking budgets, non-default values for parameters such as temperature and top_p, assistant prefill, and forcing tool use, whether a specific tool or any tool, all return a 400 error.
- Cybersecurity tasks: because the model's cyber capabilities are comparable to Opus 5, Anthropic launched it with safeguards under which higher-risk security tasks visibly fall back to Sonnet 5. On the API, a declined request returns a refusal unless the developer turns on the beta server-side fallback, available on the Claude API only. Finding and fixing bugs as part of routine development is unaffected.
The model is available on the Claude Platform, AWS, Google Cloud and Microsoft Azure, and the documentation states it will not be retired before 28 September 2027, so anyone building on it today has about a year before being forced to move.
How to measure cost per task in your business
The question worth putting to whoever runs your AI integration is not what the model costs, but what one accepted task costs us. The measurement is simple:
- Pick a real sample: 20 to 50 tasks from your actual work, such as customer replies or invoices to extract data from, not made-up examples.
- Run it on both models: your current model and Sonnet 5.5, at two effort levels at least. The API returns input and output token counts with every response, so log them alongside the run time.
- Have an employee judge quality: someone who knows the work decides which outputs are acceptable as they are.
- Compute cost per accepted task: total cost divided by the number of accepted outputs. A cheaper model that needs a third of its tasks redone is not cheaper.
- Use the pricing options: non-urgent work, such as processing an invoice archive, can run through the Batch API at half price: $1 input and $5 output per million tokens. Context that repeats in every request, such as company policies or a product catalogue, is read from the cache at $0.20 instead of $2, and the minimum cacheable prompt is now 512 tokens instead of 1,024.
The Origami view
Price tables invite comparisons by the cost of a million tokens, but this release is a clear example that the number that matters is the cost of a task. The same model at the same price can get cheaper because it works in fewer steps, and it can get more expensive if thinking switches on by default or someone raises the effort level without measuring. A business paying a monthly AI bill without knowing the cost of a single task is managing a line item it cannot see.
That is why we think any system that uses a language model inside a business should make the model name and effort level a setting that changes without reprogramming, and should log tokens per task in a monthly report. Then moving to a new release, such as Sonnet 5.5 today and Haiku 5.5, which Anthropic says is coming in the next few weeks, becomes a decision measured in days, not a new project.
The bottom line
Sonnet 5.5 did not cut the list price, but it may cut your bill if the previous model spends many steps and tokens on your kind of work. Before you switch:
- Ask your developer to review the thinking settings and the code that reads responses, then test the integration in a staging environment before production.
- Measure the cost per accepted task on a real sample, at two effort levels at least.
- If you are on Sonnet 4.6 or older, account for the extra tokens the same text now produces before assuming any saving.
- Keep Opus 5.5 for open-ended tasks that need judgment, move well-defined repetitive tasks to Sonnet 5.5, and watch the numbers for a full month.
Sources
- Anthropic — Introducing Claude Sonnet 5.5 (28 September 2026)
- Claude docs — Pricing, including the Batch API and prompt caching
- Claude docs — Models overview: context window, max output, default effort and retirement dates
- Claude docs — Migrating to Claude Sonnet 5.5
- TechCrunch — Anthropic releases Sonnet 5.5 (28 September 2026)
Frequently asked questions
Is Claude Sonnet 5.5 cheaper than Sonnet 5?+
Its price per million tokens is the same: $2 for input and $10 for output. But Anthropic says it costs up to 30% less per task in its testing, because it does the work with fewer tokens and steps. The saving varies by kind of work, and you can only know it for your business by testing a sample of your real tasks.
When should I choose Opus 5.5 over Sonnet 5.5?+
Opus 5.5 costs twice as much: $4 input and $20 output per million tokens. Anthropic says it remains clearly stronger at complex, open-ended work that needs sustained judgment, while Sonnet 5.5 is strongest at well-scoped everyday tasks, bug fixing, and producing documents and slides. Repetitive tasks with clear rules usually belong on Sonnet.
Is changing the model name in our system enough to switch?+
Not always. Systems that turn thinking off with the disabled setting will get a 400 error and must use between_tools, non-default values for parameters such as temperature are rejected, as is forcing tool use, and a response can begin with a thinking block before the text. Ask your developer to follow the migration guide and test the integration in staging first.
What effort level suits everyday business tasks?+
The API default is high. Anthropic's documentation recommends starting at medium or low for chat and latency-sensitive work, and at medium for well-specified coding and tool tasks. Try two levels on a sample of your tasks, compare the cost per accepted task, and do not assume the highest level gives the best result.
Follow Origami in Google
Pin Origami as a preferred source and our articles will surface first for you in Google Search and Top Stories.

Related articles
- Artificial IntelligenceEleven v4 Speaks Arabic in 100 Milliseconds. The Hard Part Is What It Tells Your CustomerElevenLabs' Eleven v4 and v4 Turbo bring ~100 ms voice generation and 90+ languages including Arabic. Prices, limits, and what to fix before a voice agent answers customers.
- Artificial IntelligenceXiaomi Opens MiMo-V2.6 Under MIT, and Its Light Model Reads a Million Tokens for 14 CentsXiaomi released MiMo-V2.6 open source under MIT: one model for text, audio, images and video, with a 1M-token context and prices from $0.14 per million tokens.
- Artificial IntelligenceAn Open Agent Built for Days of Work, Not Minutes: Atria Dawn and Its Real Running BillShanghai AI Lab released Atria Dawn under MIT: 744 billion parameters aimed at long multi-step tasks. What it is actually good for, and what running it in-house costs.
- Artificial IntelligenceThe Model Race Taps the Brakes: Outside Evaluators Get a Badge and a Desk Inside AnthropicOn September 12, 2026 Anthropic's CEO called for slowing AI capability gains and committed to letting independent evaluators inside. Here is what actually changes for your tech plan.
- Artificial IntelligenceSWE-2 Matches Frontier Coding at a Third of the Cost, Then Fails One TestCognition released SWE-2 on September 10, 2026: 50.0% on FrontierCode versus Fable 5.1's 50.9% at 64% lower cost, yet 28 points behind on Terminal-Bench 4.
- Artificial IntelligenceOpenAI's Two New Image Models Edit One Part of Your Product Photo Without a ReshootOpenAI shipped GPT Image 2.5 in two versions on September 8, 2026: Sunburst for editing precision, Flare for speed. What changed for your store, what it costs, where it helps.
Have a project in mind?
We build custom systems, apps and websites for your business. Tell us your idea and we will give you a straight answer on it.
