Claude Sonnet 5.5 Release Puts a 70.6% Benchmark Score Against an Unchanged Rate Card
Anthropic released Claude Sonnet 5.5 with a reported 70.6% Terminal-Bench 4.0 score, while keeping its published Sonnet token rates unchanged.
That combination creates the real story. Anthropic is not asking developers to pay more per token for its newer everyday model. It says the model also generates output over 30% faster and uses fewer tokens for many tasks.
The official launch compares that 70.6% result with 10.3% for Sonnet 5. It also places Sonnet 5.5 above the company’s reported 66.4% result for Opus 5.5 on the same benchmark.
Those numbers make the Claude Sonnet 5.5 release look like more than a routine model refresh. They pressure the traditional split between a fast default model and a premium model reserved for difficult work.
However, the headline result needs context. Benchmark configurations, reasoning effort, agent scaffolding, fallbacks, and token use can materially change both scores and operating costs.
Independent testing already presents a more complicated picture than Anthropic’s launch chart. Sonnet 5.5 appears highly competitive, but its best results do not automatically translate into the cheapest production deployment.
The Claude Sonnet 5.5 Release Changes the Default Model Calculation
Anthropic is positioning Sonnet 5.5 as the model teams can use by default, not as a limited alternative to Opus.
The model became available on September 28, 2026, across Anthropic’s apps and developer platform. Anthropic also says it is available through Amazon Web Services, Google Cloud, and Microsoft Azure.
The release targets well-scoped coding, bug fixing, document creation, presentations, spreadsheets, and everyday agent workflows. Anthropic continues to position Opus 5.5 for open-ended work requiring sustained judgment.
That distinction matters because many business workloads are closer to the Sonnet category. A support reply, code review, document revision, or structured analysis rarely requires maximum reasoning on every request.
Anthropic reports that Sonnet 5.5 produced output over 30% faster than Sonnet 5. It also says the model can reduce total task costs despite keeping the same published token rates.
The claim depends on task efficiency. A model can carry an unchanged rate card while becoming cheaper per completed task if it uses fewer tokens, tools, or retries.
Anthropic supplied several early-customer examples to support that argument. Slack reported about 14% fewer output tokens across its offline Slackbot evaluations, without changing its prompts.
Zendesk reported that support tickets were processed 20% faster in its testing. Atlassian said its Rovo agents could run up to 30% faster than with Sonnet 5.
Balyasny Asset Management tested the model on 2,441 private finance tasks. It reported substantially lower token use per answer than Sonnet 5 across analysis, extraction, forecasting, and retrieval work.
These are company-selected launch examples, not controlled comparisons across every workload. Still, they illustrate what Anthropic wants buyers to measure: completed work rather than isolated token prices.
The practical change is therefore broader than a benchmark score. Teams evaluating the model must compare latency, success rates, retries, tool calls, and review time together.
A faster model that finishes more tasks correctly can change queue capacity and user experience. It can also reduce the human effort spent correcting incomplete work.
The release includes a new model identifier, claude-sonnet-5-5. Developers migrating from Sonnet 5 must update more than that identifier in some configurations.
Anthropic’s migration guidance documents changes involving thinking settings, forced tool selection, content blocks, and computer-use tools. Some older request patterns return errors.
For example, thinking runs by default when developers omit the relevant field. Applications that assume the first returned block always contains normal text can therefore break.
The model also replaces disabled up-front thinking with a between_tools setting at supported effort levels. That behavior matters for applications designed around low-latency responses or predictable reasoning budgets.
These compatibility details complicate the idea of a cost-free upgrade. The published rates may remain unchanged, but migration work and evaluation time still carry operational costs.
That is why teams should treat Sonnet 5.5 as a new runtime, not merely a better checkpoint behind the same API contract.
A 70.6% Terminal-Bench Score Creates Pressure Above Sonnet
The surprising comparison is not Sonnet 5.5 against its predecessor. It is Sonnet 5.5 against Anthropic’s premium Opus line.
Terminal-Bench evaluates agents working through a command-line interface. The tasks require models to inspect environments, use tools, modify artifacts, and complete multistep objectives.
Version 4.0 contains 66 community-contributed and maintainer-reviewed tasks. Its categories include software, science, machine learning, operations, hardware, security, and media.
The benchmark methodology emphasizes final artifacts rather than persuasive explanations. An agent receives credit when its work passes the grader, not when its response merely sounds plausible.
Anthropic reports a 70.6% result for Sonnet 5.5 and a 10.3% result for Sonnet 5. It reports 66.4% for Opus 5.5 at that model’s highest evaluated effort.
The gap between the two Sonnet generations is unusually large. It indicates that something beyond incremental language quality changed within Anthropic’s agentic coding stack.
The result also reverses the expected product hierarchy on this specific test. A lower-cost family member reportedly outscored Anthropic’s premium model on complex terminal work.
That does not make Sonnet 5.5 universally better than Opus 5.5. Anthropic explicitly says Opus remains stronger on complex, open-ended assignments requiring sustained judgment.
Other published evaluations support that qualification. On CursorBench 4.0, Sonnet 5.5 scored 55.5%, while Opus 5.5 scored 57.8%.
On GDPval-AA v2.1, which evaluates occupational tasks, the reported scores were 1,844 for Sonnet 5.5 and 1,846 for Opus 5.5. The models were nearly level there.
FrontierCode produced another mixed result. Sonnet 5.5 reached 52.1% at one effort setting, while Opus 5.5 scored 54.4%.
Together, those results describe a narrower reversal. Sonnet 5.5 appears especially strong when a task has clear goals, usable tools, and verifiable completion conditions.
Opus retains an advantage when success depends on ambiguous judgment, broader planning, or maintaining quality through an open-ended assignment.
For developers, that split encourages model routing. A system can send routine implementation and bounded agent tasks to Sonnet, while reserving Opus for architecture or difficult escalation cases.
Anthropic’s launch materials offer one example of this division. A creator described using Opus to establish a game’s architecture, then trusting Sonnet 5.5 to implement it.
This model pairing is more important than a simple leaderboard victory. It suggests that premium reasoning and high-volume execution may become separate stages within one workflow.
The same pattern fits document and knowledge work. A premium model might define an analysis plan, while Sonnet handles extraction, drafting, revisions, and formatting.
Teams already building an engineering knowledge base can test this structure against repository documentation and review records. Their own accepted outputs matter more than a generic ranking.
If Sonnet 5.5 consistently handles the execution stage, Opus faces pressure from inside Anthropic’s own product family. Developers will ask why every difficult-looking task needs the premium model.
That question becomes especially significant when the cheaper model also responds faster. Latency often determines whether users tolerate an agent in an interactive coding loop.
The Benchmark Leap Reflects a Better Agent Loop, Not Just Better Answers
The Claude Sonnet 5.5 benchmark result points toward more efficient tool use, but Anthropic has not isolated one cause for the full increase.
Agentic benchmarks measure a combined system. The underlying model matters, but so do prompts, tools, reasoning effort, context management, time limits, and fallback behavior.
Anthropic says early testers observed fewer steps and more batched tool calls. Lovable reported roughly half as many shell runs and about one-third fewer tool calls in its internal coding evaluations.
CodeRabbit also reported that Sonnet 5.5 used fewer output tokens and showed better judgment across tasks with different complexity levels. It noted less unnecessary web searching than with Sonnet 5.
These observations offer a plausible mechanism for the Terminal-Bench improvement. An agent that explores less aimlessly can preserve time and context for actions that change the final artifact.
Tool efficiency also affects reliability. Every shell command, browser action, or external request creates another opportunity for failure, latency, or malformed output.
A model that chooses a shorter valid path can therefore improve completion rates without producing dramatically better prose. Terminal work rewards this type of discipline.
Anthropic added five effort levels for Sonnet 5.5. The setting controls how long the model reasons and checks its work before or between actions.
Higher effort can improve difficult tasks, but it also consumes more tokens and time. Anthropic recommends that teams evaluate several settings instead of transferring old assumptions from Sonnet 5.
That recommendation is easy to overlook. A benchmark’s best result usually reflects a deliberate configuration, while production systems often use a default or cost-controlled setting.
The launch chart reports the 70.6% figure within Anthropic’s evaluation framework. Independent evaluators can obtain different results when they change the harness or effort level.
Artificial Analysis, for example, reported 64% on its own Terminal-Bench 4.0 run. Its independent evaluation placed Sonnet 5.5 among the leading models, but highlighted heavy token use at maximum effort.
It found that Sonnet 5.5 consumed more output tokens per Intelligence Index task than any model it had measured. That result challenges a simple efficiency narrative.
There is no necessary contradiction between the two findings. Sonnet 5.5 can be efficient at lower settings and token-intensive when pushed toward its maximum measured capability.
The distinction between rate and total consumption is crucial. An unchanged rate card does not guarantee an unchanged bill when the model reasons longer.
Anthropic says low or medium effort can beat Sonnet 5’s best scores at a fraction of the completed-task cost on several evaluations. Independent testing suggests maximum effort has a different profile.
Production buyers should therefore evaluate a curve, not one point. The useful comparison plots task success against latency, tokens, retries, and human review.
A coding team could start with a representative set of repository tasks. Those tasks should include bug fixes, refactors, test creation, dependency changes, and unfamiliar code navigation.
Each run should use the same environment and acceptance checks. Reviewers should record whether the patch works, stays within scope, and requires human correction.
The test should also count failed tool calls and elapsed time. Those measurements reveal whether a higher headline score translates into a better development loop.
Teams should repeat the exercise at multiple effort levels. If medium effort completes most routine work, maximum effort may add cost without providing enough additional value.
The best configuration may vary within one product. A fast interactive assistant needs different settings from an overnight migration agent with extensive validation.
This is the central mechanism behind the release. Anthropic is giving developers more control over how much computation Sonnet spends, while claiming better outcomes across that range.
What the Numbers Do Not Settle
The benchmark headline is credible as a reported result, but it cannot establish production reliability or universal cost savings by itself.
The first limitation is configuration sensitivity. Anthropic’s 70.6% score and Artificial Analysis’s 64% score both describe Sonnet 5.5, yet they come from different evaluation setups.
The second limitation is fallback behavior. Some evaluation systems can route a declined or unsupported request to another model under defined conditions.
Artificial Analysis observed fallback on a small fraction of its tasks. Vals also documents provider-side fallback as a factor that can affect leaderboard interpretation.
Fallback is not inherently improper. It can represent the actual product behavior customers receive, especially when providers use routing to maintain safety or availability.
However, a fallback-assisted result answers a different question from a pure-model result. Buyers should know whether they are evaluating a model, a provider gateway, or an entire managed agent.
The third limitation concerns benchmark saturation. A 70.6% score leaves meaningful room for failure, yet it also narrows the benchmark’s ability to separate future models.
When leading systems complete most tasks, difficult edge cases matter more. Small prompt or harness changes can also move rankings without transforming normal user experience.
Terminal-Bench remains useful because its tasks require real action and produce checkable artifacts. Still, no single benchmark represents every codebase, toolchain, security policy, or approval process.
The fourth limitation is total resource use. Artificial Analysis found that Sonnet 5.5’s maximum-effort configuration used about 193,000 output tokens per Intelligence Index task.
That measurement does not describe every request. It does show why teams should not infer completed-task cost from the published token rate alone.
At maximum effort, Artificial Analysis found Sonnet 5.5 outside the most efficient frontier in its comparison. Other configurations offered different balances.
The fifth limitation involves safety behavior. Anthropic says Sonnet 5.5 is its first Sonnet model to launch with cyber safeguards similar to those used for its most capable models.
Higher-risk cyber requests can fall back to Sonnet 5. The model also includes classifiers intended to stop attempts to extract its reasoning.
These protections respond to stronger capabilities, but they can create new refusal patterns. A legitimate security workflow might behave differently after migration.
The cyber safeguards therefore represent both a safety measure and an operational variable. Security teams need evaluation cases that cover authorized defensive work.
The sixth limitation is launch-partner selection. Anthropic’s customer testimonials provide concrete data, but the company chose which examples appeared in its announcement.
Slack, Zendesk, Box, Lovable, Atlassian, and other partners tested workloads that matter to them. Their outcomes do not establish the same gains for unrelated applications.
A finance retrieval system, coding agent, and customer-support workflow place different demands on a model. They also apply different standards for acceptable errors.
Teams should reproduce the claimed improvements with their own data and graders. A model that saves tokens but increases review time has not improved the total workflow.
The reverse can also happen. A model that uses more tokens may still be economical if it prevents failures, reduces retries, or completes work that previously required escalation.
That is why the strongest interpretation remains conditional. Sonnet 5.5 appears to move the capability-cost boundary, especially for bounded agent work.
The release does not eliminate the need for Opus, custom evaluations, or human review. It makes the decision about when to use each one more consequential.
Three Signals Will Show Whether Sonnet 5.5 Changes the Market
The next test is whether developers reproduce Anthropic’s results at ordinary effort levels and move real workloads away from premium models.
The first signal is independent benchmark replication. Evaluators should publish results with effort settings, harness details, fallback counts, token consumption, and task-level failures.
A replicated result near Anthropic’s score would strengthen the claim that Sonnet 5.5 represents a major agentic improvement. Large variation would make configuration the more important story.
The difference between 70.6% and 64% already shows why disclosure matters. Both scores indicate strong performance, but they imply different comparisons with competing systems.
The second signal is production routing. Watch whether coding tools and enterprise platforms make Sonnet 5.5 their default model for routine agents.
Default placement matters more than optional availability. It reveals whether vendors trust the model’s latency, reliability, refusal behavior, and completed-task economics.
A shift from Opus to Sonnet for implementation work would support Anthropic’s product strategy. Limited adoption would suggest that premium reasoning still provides essential reliability.
Early testimonials point toward routing rather than wholesale replacement. CodeRabbit plans to move simpler and moderate reviews first, then expand based on results.
That approach is sensible. It treats model selection as an operational policy instead of a brand preference.
The third signal is completed-task cost across effort levels. Buyers should look for measurements that include output tokens, tool calls, retries, latency, and reviewer intervention.
If medium effort preserves most of the benchmark gain, Sonnet 5.5 strengthens its case as a high-volume default. If maximum effort is routinely required, the economic advantage becomes narrower.
Developers must also monitor migration errors. The new thinking behavior, tool-choice rules, safety fallbacks, and content-block handling can affect existing integrations.
Anthropic’s documentation advises teams to rerun their effort sweeps and rebaseline costs. That instruction is more important than the unchanged rate card.
The Claude Sonnet 5.5 release ultimately challenges a familiar assumption: the premium model is always the safer choice for serious agent work.
Anthropic’s own results show Sonnet leading Opus on one important terminal benchmark. Other tests still favor Opus, particularly where sustained judgment matters.
That creates a clearer division of labor. Sonnet 5.5 can handle fast, bounded execution, while Opus remains the escalation path for ambiguous decisions.
The market impact will depend on whether that division survives contact with real repositories, documents, support queues, and security controls.
Teams evaluating Claude Sonnet 5.5 should begin with completed tasks, not isolated prompts. Build a fixed test set, run several effort levels, and record every retry.
Compare the model with both Sonnet 5 and the premium alternative already used in production. Include migration work, reviewer time, refusals, and failed tool calls.
Then ask the decision that matters: does Claude Sonnet 5.5 finish enough real work, with enough reliability, to become your new default?



