NEWS

Anthropic launches Sonnet 5.5, and the leap in Terminal Bench impresses

Anthropic launched Claude Sonnet 5.5 on September 28. This is the second model in the Claude 5.5 family. The proposal addresses well-defined tasks.

Anthropic launches Sonnet 5.5, and the leap in Terminal Bench impresses
Image: Redação iMasters

Anthropic launched Claude Sonnet 5.5 on September 28. This is the second model in the Claude 5.5 family. The proposal addresses well-defined tasks, such as fixing code, creating documents, presentations, and spreadsheets.

The gains show up on two fronts. Also, the model responds more than 30% faster and can cost up to 30% less per task.

Note that the pricing table stayed the same. The savings come from lower token consumption per job.

Anthropic kept the per-token price of Sonnet 5

However, the prices remain the same as the previous generation. They are US$2 per million input tokens and US$10 per million output tokens.

Cache reads cost US$0.20 per million. Therefore, those who already used Sonnet 5 keep the same calculation basis.

The real reduction happens elsewhere. Since the model needs fewer tokens to complete each task, the final bill drops.

The Terminal Bench number calls for a second look

Here is the most striking result of the announcement. On Terminal Bench 4.0, Sonnet 5.5 reached 70.6%.

However, Sonnet 5, for comparison, stood at 10.3%. Also, this test measures complex tasks executed on the command line.

The other numbers also rose. On CursorBench 4.0, the new model scored 55.5% against 34.1% for the previous version.

On FrontierCode 1.1, the gain was smaller. The index went from 42.4% to 46.2%.

Also, outside of coding, two results stand out. GDPval AA v2.1 jumped from 1,449 to 1,844 points, while OSWorld 2.1 rose from 57% to 80.1%.

Anthropic keeps Opus 5.5 for open-ended work

The company made the division clear. Opus 5.5 remains stronger for complex work that requires prolonged judgment.

However, there is an interesting caveat. In some evaluations, Sonnet 5.5 came close to the higher-tier model when operating at maximum effort.

This detail matters when choosing. Not every task needs the most expensive model in the catalog.

What testers reported in practice

The improvement in programming showed up in real-world use. According to the company, the model understands codebases quickly and solves tasks in fewer steps.

Tool-use efficiency was also highlighted. Daniel Vogel, chief operating officer at Epic Games, said the model reached the quality level expected from a higher tier.

On Slack's side, the account was direct. Curtis Allen, principal engineer, said the model outperformed Sonnet 5 on nearly all internal Slackbot evaluations, without any prompt changes.

There is also an internal creation test. The model received quarterly materials, transcripts from a public company, and a presentation template. The result was a ten-slide operational review that two specialists considered ready to send.

Safeguards arrive at the Sonnet line

This is the first time a Sonnet model has shipped with this package. It brings cybersecurity safeguards and fallback mechanisms similar to those in Opus 5.5.

How this works deserves attention from those working in security. Also, routine development tasks remain available as usual.

Requests considered riskier, on the other hand, can be routed to Sonnet 5. In biology, the protections remain the same as in the previous version.

Where to use it and what to measure before migrating: Anthropic

The model is already available on every platform. The list includes Amazon Web Services, Google Cloud, and Microsoft Azure.

Before switching, measure the cost per completed task. Price per token matters little when consumption changes this much between versions.

Then, test it on your own repository. Public benchmarks indicate direction, while your own code decides the choice.

However, it's also worth reviewing the routing between models. Simple tasks on Sonnet and open-ended work on Opus tend to perform better than using just one model for everything.

Finally, watch the fallback behavior. Security flows may fall back to Sonnet 5, and that changes the expected result.

Claude Haiku 5.5 arrives in the coming weeks to complete the family.

Follow our profile on Instagram!

Translated from the Brazilian Portuguese original · Read the original

More from Redação iMasters
View profile →
Read also