Back to Blog
Artificial Intelligence

Sonnet 5.5 Costs the Same per Token as Sonnet 5, and Up to 30% Less per Task

Origami TeamEditorial Team
6 min read
Sonnet 5.5 Costs the Same per Token as Sonnet 5, and Up to 30% Less per Task
Like what we publish? Pin Origami as a preferred source on Google.Add as a preferred source on Google

Sonnet 5.5 Costs the Same per Token as Sonnet 5, and Up to 30% Less per Task

On 28 September 2026 Anthropic released Claude Sonnet 5.5, the second model in the 5.5 generation after Opus 5.5, which shipped on 22 September. The price did not move from Sonnet 5: $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens read from the prompt cache. A token is the unit of text AI models are billed by.

Even so, Anthropic says the new model costs up to 30% less per task in its testing, and writes its output more than 30% faster. The reason is not a discount. The model simply needs fewer tokens to do the same work. That is the point worth any business's attention if it pays an AI bill: the price per token does not tell you what a task costs.

Why the bill falls while the token price stays put

A model bill is the number of tokens multiplied by their price. And the token count is not set by the length of your question alone. It depends on what the model does to reach an answer: how many steps it takes, how often it calls a tool or searches, and how much it thinks before it writes. Anthropic's documentation states that thinking tokens are billed as output tokens, the more expensive kind.

Alongside the announcement, Anthropic published results from companies that tested the model before launch:

  • Balyasny Asset Management: on 2,441 finance tasks covering Q&A, extraction, analysis and forecasting, Sonnet 5.5 used about 121,000 tokens per answer against 497,000 for Sonnet 5, and scored higher.
  • Slack: it beat Sonnet 5 on almost all of Slack's offline Slackbot evaluations, in fewer steps and with about 14% fewer output tokens, without any change to the prompts.
  • Lovable: in its coding evaluations it needed a third fewer tool calls and roughly half the shell runs to finish a task.
  • Box: it was more accurate than the previous model, 2.4 times faster, and used 12% fewer total tokens.

Notice the range: from 12% fewer tokens to roughly three quarters fewer, depending on the kind of work. These are the companies' own tests as published by Anthropic, not independent measurements. In practice, the real saving in your business can only be known by trying it on your own tasks.

Sonnet 5.5 or Opus 5.5: the difference in price and capability

Opus 5.5 charges twice as much for input and output, while cache reads cost the same, and on most published benchmarks the gap is small. But Anthropic itself gives each model a different job.

ItemSonnet 5.5Opus 5.5
API model IDclaude-sonnet-5-5claude-opus-5-5
Price (input / output per million tokens)$2 / $10$4 / $20
Cache write (5 minutes)$2.50$5
Cache read$0.20$0.20
Batch API (input / output)$1 / $5$2 / $10
Context window / max output1M tokens / 128K tokens1M tokens / 128K tokens
Default effort on the APIhighmedium
Terminal-Bench 4.0 — terminal coding70.6%66.4%
CursorBench 4.0 — tasks from real coding sessions55.5%57.8%
OSWorld 2.1 — computer use80.1%81.8%
GDPval-AA — professional work across 44 occupations (rating)18441846
Humanity's Last Exam — with tools64.5%67.7%

Figures come from Anthropic's announcement and its pricing and models documentation. The Opus 5.5 Terminal-Bench score is its highest, at Xhigh effort. For comparison, Sonnet 5 rated about 1449 on GDPval-AA.

The numbers are close, but the same announcement says benchmarks capture only one facet of capability, and that in Anthropic's own testing and that of external testers Opus 5.5 remains clearly stronger at complex, open-ended work that needs sustained judgment. Sonnet 5.5 is at its best on well-scoped everyday tasks, bug fixing, and producing documents, slides and spreadsheets.

Our reading, in operational terms: answering a customer query under clear rules, extracting invoice data, summarising a meeting, or drafting a deck from a fixed template are Sonnet jobs. Designing a new system, an open-ended analysis that weighs options, or a decision that rests on long and tangled context are where Opus may earn its price. We compared Opus 5.5 with the top-tier Fable 5.1 in detail in an earlier article.

Effort: the setting that moves the bill more than the model name

Sonnet 5.5 runs at five effort levels: low, medium, high, xhigh and max. The higher the level, the longer the model works and the more it checks itself, so the cost per task rises and the score usually does too. The default is medium inside the Claude apps and Claude Code, and high when you call it through the API.

Anthropic says that on several benchmarks Sonnet 5.5 at low or medium effort beat Sonnet 5's best score for about a tenth of the cost per task. On FrontierCode, which measures whether a code change could be merged without human edits, it scored 10 points above Sonnet 5 at the same high setting, for about one fifteenth of the cost per task.

One note in the announcement deserves a pause: on FrontierCode, Sonnet 5.5 scored lower at max effort than at xhigh, because it more often ran Claude Code's code-review skill, which splits the review across many subagents; in two cases examined by Cognition, this led to a timeout or to edits beyond the task. The highest effort is not always the best result, but it usually costs the most.

The migration guide also warns that the levels were recalibrated, so a given level does not produce the same amount of thinking as it did on Sonnet 5. It recommends starting at high for general work, at medium for well-specified coding and multi-step tool tasks, and at medium or low for chat and anything latency-sensitive. In other words, a system that calls the model through the API without setting a level will run at high, and may pay more than its routine tasks need.

Before your developer changes the model name

Switching models looks like a one-line change, but the migration guide lists changes that can stop a working integration or alter its bill:

  • Turning thinking off has changed: systems that disable thinking on Sonnet 5 with thinking: {"type": "disabled"} will get a 400 error from Sonnet 5.5. The replacement is the new between_tools setting, which is accepted only at low, medium and high effort.
  • Thinking is on by default: a request that sets no thinking field runs with adaptive thinking on Sonnet 5.5. Anyone moving from Sonnet 4.6 or earlier, or from Haiku 4.5, where thinking was off by default, will now pay for thinking tokens on requests that never set a thinking field.
  • The same text, more tokens, if you come from generation 4: Sonnet 5.5 uses Sonnet 5's tokenizer, and the documentation says the same text produces about 30% more tokens than on Sonnet 4.6, Sonnet 4.5 and Haiku 4.5, depending on the content. The per-task saving Anthropic announces is measured against Sonnet 5, not against generation 4.
  • Response shape: a response can begin with a thinking block before the text, so code that reads the first element of the reply as the text will break.
  • Rejected settings: fixed thinking budgets, non-default values for parameters such as temperature and top_p, assistant prefill, and forcing tool use, whether a specific tool or any tool, all return a 400 error.
  • Cybersecurity tasks: because the model's cyber capabilities are comparable to Opus 5, Anthropic launched it with safeguards under which higher-risk security tasks visibly fall back to Sonnet 5. On the API, a declined request returns a refusal unless the developer turns on the beta server-side fallback, available on the Claude API only. Finding and fixing bugs as part of routine development is unaffected.

The model is available on the Claude Platform, AWS, Google Cloud and Microsoft Azure, and the documentation states it will not be retired before 28 September 2027, so anyone building on it today has about a year before being forced to move.

How to measure cost per task in your business

The question worth putting to whoever runs your AI integration is not what the model costs, but what one accepted task costs us. The measurement is simple:

  • Pick a real sample: 20 to 50 tasks from your actual work, such as customer replies or invoices to extract data from, not made-up examples.
  • Run it on both models: your current model and Sonnet 5.5, at two effort levels at least. The API returns input and output token counts with every response, so log them alongside the run time.
  • Have an employee judge quality: someone who knows the work decides which outputs are acceptable as they are.
  • Compute cost per accepted task: total cost divided by the number of accepted outputs. A cheaper model that needs a third of its tasks redone is not cheaper.
  • Use the pricing options: non-urgent work, such as processing an invoice archive, can run through the Batch API at half price: $1 input and $5 output per million tokens. Context that repeats in every request, such as company policies or a product catalogue, is read from the cache at $0.20 instead of $2, and the minimum cacheable prompt is now 512 tokens instead of 1,024.

The Origami view

Price tables invite comparisons by the cost of a million tokens, but this release is a clear example that the number that matters is the cost of a task. The same model at the same price can get cheaper because it works in fewer steps, and it can get more expensive if thinking switches on by default or someone raises the effort level without measuring. A business paying a monthly AI bill without knowing the cost of a single task is managing a line item it cannot see.

That is why we think any system that uses a language model inside a business should make the model name and effort level a setting that changes without reprogramming, and should log tokens per task in a monthly report. Then moving to a new release, such as Sonnet 5.5 today and Haiku 5.5, which Anthropic says is coming in the next few weeks, becomes a decision measured in days, not a new project.

The bottom line

Sonnet 5.5 did not cut the list price, but it may cut your bill if the previous model spends many steps and tokens on your kind of work. Before you switch:

  • Ask your developer to review the thinking settings and the code that reads responses, then test the integration in a staging environment before production.
  • Measure the cost per accepted task on a real sample, at two effort levels at least.
  • If you are on Sonnet 4.6 or older, account for the extra tokens the same text now produces before assuming any saving.
  • Keep Opus 5.5 for open-ended tasks that need judgment, move well-defined repetitive tasks to Sonnet 5.5, and watch the numbers for a full month.

Sources

#Claude Sonnet 5.5#Anthropic#AI Costs#Language Models#Business Automation

Frequently asked questions

Is Claude Sonnet 5.5 cheaper than Sonnet 5?+

Its price per million tokens is the same: $2 for input and $10 for output. But Anthropic says it costs up to 30% less per task in its testing, because it does the work with fewer tokens and steps. The saving varies by kind of work, and you can only know it for your business by testing a sample of your real tasks.

When should I choose Opus 5.5 over Sonnet 5.5?+

Opus 5.5 costs twice as much: $4 input and $20 output per million tokens. Anthropic says it remains clearly stronger at complex, open-ended work that needs sustained judgment, while Sonnet 5.5 is strongest at well-scoped everyday tasks, bug fixing, and producing documents and slides. Repetitive tasks with clear rules usually belong on Sonnet.

Is changing the model name in our system enough to switch?+

Not always. Systems that turn thinking off with the disabled setting will get a 400 error and must use between_tools, non-default values for parameters such as temperature are rejected, as is forcing tool use, and a response can begin with a thinking block before the text. Ask your developer to follow the migration guide and test the integration in staging first.

What effort level suits everyday business tasks?+

The API default is high. Anthropic's documentation recommends starting at medium or low for chat and latency-sensitive work, and at medium for well-specified coding and tool tasks. Try two levels on a sample of your tasks, compare the cost per accepted task, and do not assume the highest level gives the best result.

Follow Origami in Google

Pin Origami as a preferred source and our articles will surface first for you in Google Search and Top Stories.

Add as a preferred source on Google

Related articles

Have a project in mind?

We build custom systems, apps and websites for your business. Tell us your idea and we will give you a straight answer on it.

One session. Twenty minutes. No commitments.