AI Affairs, home

Thursday 1 October 2026

Technology

Anthropic releases Sonnet 5.5 with faster output and plans Haiku 5.5

Anthropic says fewer tokens and tool calls can cut costs per task by up to 30%, with API rates of $2 per million input tokens and $10 per million output tokens.

Detailed view of server rack units with glowing status lights and technology.
Photo: panumas nikhomkhai via Pexels

Anthropic released Claude Sonnet 5.5 on 28 September, TechCrunch reported, with the company claiming faster output and a lower cost to complete tasks than with Sonnet 5. The release follows Opus 5.5, which Anthropic introduced less than a week earlier.

Key points

  • Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens through the API.
  • Anthropic says output is more than 30% faster and total cost per task can fall by as much as 30%.
  • On Anthropic’s reported Terminal-Bench 4.0 test, Sonnet 5.5 scored 70.6%, against 66.4% for Opus 5.5.
  • Anthropic says the model will receive cyber safeguards used for Fable and Opus, and plans to release Haiku 5.5 in the coming weeks.

Sonnet 5.5 keeps its token prices

Anthropic prices Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens through its API, VentureBeat reported. Opus 5.5 costs $4 and $20 respectively. As AI Affairs reported on the Opus release, Anthropic also lowered prices for that model and published stronger coding scores.

Anthropic puts the gain in Sonnet 5.5’s output speed over Sonnet 5 at more than 30%, and says the cost of finishing a task can fall by as much as 30%. Its explanation centres on the amount of work the model does at the same token price: it uses fewer tokens and makes fewer tool calls, or requests to other software, while working towards an answer. A price per token is like a price per sheet of paper. The final bill also depends on how many sheets a job uses.

Preparing a presentation could take less time and cost less to generate if the task needed fewer model-generated tokens and fewer tool calls, even though the charge for each token stayed the same.

Anthropic says the model is suited to defined work such as debugging software and producing documents, presentations and spreadsheets. The company describes Opus 5.5 as its choice for more ambiguous work requiring sustained judgment. That distinction matters to companies choosing between the models: a cheaper completed task depends on the work the model must actually finish, as well as on the rate charged for its tokens.

Anthropic also offers settings that change how much effort the model spends on a task. It says Sonnet 5.5 at Low or Medium effort exceeded Sonnet 5’s best score on several evaluations at roughly one-tenth the cost per task. Claude Code and Anthropic’s consumer applications default to Medium effort, while the Claude Platform defaults to High. Lower effort can reduce time and token use; higher effort gives the model more room to check and revise its answer.

Terminal-Bench 4.0 favours Sonnet 5.5

On Terminal-Bench 4.0, an evaluation of coding tasks carried out by an agent, Anthropic reported a score of 70.6% for Sonnet 5.5, compared with 66.4% for Opus 5.5 and 10.3% for Sonnet 5 under the reported settings, VentureBeat reported. The result puts Sonnet ahead of Opus on that particular test, although Anthropic still describes Opus as better suited to difficult, open-ended work.

The pattern varies across tests. Anthropic reported 80.1% for Sonnet 5.5 on its partial OSWorld 2.1 computer-use evaluation, against 81.8% for Opus 5.5 and 57% for Sonnet 5. On GDPval-AA, which draws tasks from real-world occupations, the company reported scores of 1844 for Sonnet 5.5, 1846 for Opus 5.5 and 1449 for Sonnet 5. These are results from specified evaluations, rather than measurements of every task those models might be given at work.

Early customer tests supplied to VentureBeat by Anthropic put some of the efficiency claim into narrower working settings. Zendesk tested hundreds of support cases involving replies and escalation decisions; its director of AI, Abhinay Kathuria, said tickets were processed 20% faster and the model made fewer incorrect decisions than Claude models Zendesk uses in production. Lovable co-founder and chief technology officer Fabian Hedin said its coding evaluations found about a third fewer calls to tools and about half the command-line runs.

Those customer results concern particular support cases and coding jobs, while Anthropic’s published scores concern benchmark tasks. For a company building an application, fewer tool calls could affect both running cost and waiting time, but the reported savings are tied to the tasks and settings in which they were measured.

Anthropic applies Opus cyber safeguards to Sonnet

Anthropic says Sonnet 5.5 has cybersecurity capabilities comparable to Opus 5, making it the first Sonnet model subject to the cyber safeguards used for Fable and Opus, TechCrunch reported. Cyber capability concerns what a model can do with security-related tasks, so a model offered for routine work can still warrant the protections Anthropic applies to more capable members of its range.

Anthropic says Sonnet 5.5 does not advance the frontier of its models’ capabilities. Most of its alignment testing therefore focused on a targeted set of risks that can arise at any capability level, CNBC reported. The company also described its cyber capabilities as a large improvement over Sonnet 5’s. That combination places the model in a different position on general capability and cybersecurity, according to Anthropic’s own assessments.

The release is Anthropic’s second since chief executive Dario Amodei called earlier in September for AI companies to slow the pace at which they improve their most advanced models, CNBC reported. Sonnet 5.5 is available through Amazon Web Services, Google Cloud and Microsoft Azure, as well as Anthropic’s own platforms.

Anthropic plans to release Haiku 5.5, the smallest model in the range, in the coming weeks without a firm date, TechCrunch reported.

Topics: Foundation models, Inference