top of page

Claude Opus 5.5 Cost per Task Is Lower, but Token Prices Explain Only Half

Sep 26
12 min read

Anthropic cut Claude Opus 5.5 token rates by 20%, yet the estimated Claude Opus 5.5 cost per task can fall by more. Cache reads cost 60% less than on Opus 5, which changes the economics of long Claude Code sessions.

Claude Devs highlighted the difference through a task-cost analysis and interactive calculator published on September 22, 2026. The analysis asks developers to measure completed work, rather than compare isolated token rates.

That distinction creates the real contest between Opus 5.5 and Opus 5. A cheaper token does not guarantee a cheaper feature, migration, or debugging session. Turns, cache behavior, reasoning output, retries, and model settings determine the final result.

What Changed in Claude Opus 5.5 Cost per Task

Anthropic lowered every major token category, but cache reads received the deepest reduction.

Input and output token rates are each 20% below their Opus 5 equivalents. Cache reads are 60% lower, according to Anthropic’s model announcement.

That difference matters because Claude Code repeatedly sends earlier conversation material back to the model. The reused material often includes instructions, source files, tool results, and the work completed so far.

Prompt caching allows the service to recognize previously processed content. A cache read retrieves that reusable context at a lower rate than processing it as fresh input.

Anthropic says cache reads represent most token volume in many coding and agentic workloads. That claim describes token composition, not necessarily the largest charge on every invoice.

Output can still dominate the final cost because reasoning and generated text are billed as output tokens. A session with modest input but extensive reasoning may gain less from cheaper cache reads.

The official task-cost analysis separates those effects. It first compares both models using identical token counts, isolating the rate change.

Its illustrative session contains substantial cached context, some fresh input, and a smaller quantity of output. Under those fixed assumptions, Opus 5.5 costs about 31% less than Opus 5.

That result sits between the headline reductions. It exceeds 20% because the session benefits from cheaper cache reads, but it remains below 60% because other token categories matter.

Anthropic separately estimates that typical workloads cost about 40% less at default settings. That broader estimate includes both lower rates and the company’s expectation that Opus 5.5 completes work more efficiently.

These are different claims. The 20% and 60% figures come directly from published rate changes. The 31% example depends on an illustrative token mix.

The estimated 40% reduction adds assumptions about model behavior. Developers should not apply it automatically to every repository, prompt, or coding workflow.

Opus 5.5 became available on September 22 through the Claude API and several major cloud platforms. The model also entered Claude Code and Anthropic’s subscription products.

Its model overview lists a one-million-token context window and a default medium effort setting. Adaptive thinking is always active.

Those details affect costs beyond the new rate card. A larger available context can support longer sessions, while adaptive thinking adds billable output based on task difficulty.

The result is a lower unit cost paired with a workload-dependent total. The rate change creates the opportunity, but the agent’s path through a task decides how much of it appears.

Why a Claude Code Task Costs More Than Its Final Context

Claude Code pays for repeated processing across turns, not only the conversation visible at the end.

A Claude Code task works as a loop. The model reads context, selects a tool, examines the result, updates its reasoning, and repeats those steps.

Each loop creates another request. That request includes much of the conversation accumulated during earlier turns.

Consider a session that begins with a moderate context and grows as Claude reads files, runs tests, and receives terminal output. Its final context size does not equal its total processed input.

If the task takes many turns, the model repeatedly encounters earlier material. Prompt caching makes those repeated reads cheaper, but it does not make them free.

This mechanism explains why two sessions ending with similar code changes can have different costs. One model may locate the relevant files immediately and finish after a short validation loop.

Another may inspect the wrong subsystem, attempt a fix, encounter a failure, and retrace its path. The second session pays for more tool calls, more reasoning, and more repeated context.

The number of turns therefore acts as a cost multiplier. Each unnecessary turn carries both its new content and the conversation already accumulated.

Anthropic’s example starts with a context that grows sixfold and continues for 40 turns. The total processed input becomes far larger than the final context window.

Reducing that example to 25 turns lowers total input substantially. The saving comes from avoiding repeated passes over the same growing conversation.

This is why a reliable test command can reduce Claude Code task cost. The model receives a direct signal about whether its change works.

Without that signal, it may inspect more files or reason through several speculative explanations. A build, unit test, or reproduction script can shorten that search.

Tool batching can produce a similar effect. Reading several related files in one round may avoid additional request cycles, although indiscriminate retrieval can inflate context.

The useful target is not the fewest possible tokens. It is the shortest reliable path to a correct, verified result.

That distinction matters when comparing Opus 5.5 vs Opus 5. A newer model may generate more reasoning during one turn but require fewer turns overall.

The opposite can also occur. Opus 5.5 always uses adaptive thinking, and Anthropic says it can think more at the same named effort level.

A developer who compares only output from one request may miss the complete task pattern. The meaningful unit includes exploration, edits, tests, corrections, and final reporting.

Retries deserve particular attention. A lower-effort run that fails and must be repeated may cost more than one successful run at a higher setting.

The same applies to model downgrades. A smaller model can save tokens during a lookup, but a mistaken result may send the primary agent down an expensive detour.

Anthropic’s analysis frames this as cost per completed task. That measure rewards accurate completion and penalizes false starts, even when the underlying token rate looks attractive.

For engineering teams, the lesson is practical. Count the entire loop from a scoped request through a verified outcome, not one response or one context snapshot.

Cache Reads Create the Biggest Pricing Reversal

The deepest Opus 5.5 reduction applies to the token category that long agent sessions use most heavily.

Opus 5 charged cache reads at one-tenth of its standard input rate. Opus 5.5 lowers that relationship to one-twentieth.

Combined with the lower input rate, this produces the 60% cache-read reduction. Fresh input and output receive the smaller 20% reduction.

A high cache share therefore moves a task toward the larger saving. A short request with little reused context remains closer to the basic token-rate reduction.

Anthropic’s calculator lets readers change total input, the cached portion, output, daily task volume, and an efficiency assumption. The final control represents fewer tokens used by Opus 5.5.

Leaving that efficiency assumption at zero isolates pricing. Any added reduction represents a hypothesis about how model behavior changes the task.

This separation is important. A rate card is externally verifiable, while a model’s efficiency depends on the repository and work requested.

Cache performance also depends on session behavior. Stable prompt prefixes and continuous work help the service reuse previously processed content.

Several actions can disrupt that pattern. Switching models makes the first request on the new model process the conversation under a different cache.

Changing certain settings through a cloud provider or gateway may also reduce reuse. Connecting a new tool server during a session can alter the prompt structure.

Long pauses can allow cached material to expire. The exact effect depends on the cache duration and the way requests are routed.

Cache writes add another qualification. Writing new material into the cache costs more than reading it later.

The calculator intentionally excludes cache writes from its simplified comparison. That makes the tool useful for understanding the main variables, but not a complete invoice simulator.

The first pass over a large repository can therefore remain expensive. Savings accumulate when later turns reuse what the model already processed.

Compaction introduces another tradeoff. It replaces older conversation material with a shorter summary, reducing the context resent on later requests.

However, compaction also creates a new prompt state. The immediate request must process that summary, and some detailed context may need to be retrieved again.

Clearing a session between unrelated tasks can prevent an old context from following work that no longer needs it. Clearing during one coherent task can discard useful cached context.

A model switch creates a similar boundary. Anthropic advises switching at a natural break, when the cost of rebuilding context is less likely to erase the model advantage.

Subagents further complicate the receipt. Each subagent owns a separate context window and returns a summary to the main conversation.

That separation can keep bulky file searches outside the primary context. Yet every subagent still consumes tokens and inherits a model unless configured otherwise.

The cache reduction rewards long, coherent sessions, but it does not make endless conversations optimal. Old instructions and irrelevant tool results can increase every later request.

Teams should examine both cache share and total input. A high cache rate is helpful, but an oversized conversation may still process too much material.

A well-managed session keeps reusable context warm while removing unrelated work. That balance matters more under agentic workflows than under single-response chat.

Opus 5.5 vs Opus 5 Is a Workload Test

Anthropic’s estimated savings remain a vendor projection until teams reproduce them on their own tasks.

The company says Opus 5.5 requires less compute to serve and generates output more than 30% faster than Opus 5. It also reports stronger results on several internal benchmarks.

Those findings support the case for improved cost efficiency. They do not establish a universal reduction for production codebases.

Benchmarks provide controlled comparisons, while real repositories contain incomplete tests, unusual dependencies, internal conventions, and shifting requirements. Those factors change an agent’s path.

The most uncertain variable is the number of tokens needed to finish equivalent work. Opus 5.5 may avoid false starts, but adaptive thinking can increase output on some prompts.

Its default effort is medium, while Opus 5 defaulted to high. A comparison that accepts both defaults changes more than the model version.

The migration guidance explicitly recommends recalibrating effort. Carrying an old setting forward can produce misleading results.

The model also introduces behavioral and integration changes. Thinking cannot be disabled, and several tool-use patterns require updates.

Applications using an earlier computer-use interface on the Claude API or Google Cloud must migrate to the newer toolset. Some forced tool-choice configurations now return errors.

Progress text between tool calls can also arrive through thinking blocks. An interface that does not handle those blocks may appear silent during work.

These changes are not merely migration details. Failed requests, broken tools, or missing progress displays can create retries and raise the operational cost of adoption.

A fair Opus 5.5 vs Opus 5 test should therefore hold the task constant while recording the settings. Both runs need the same repository state, acceptance criteria, and validation command.

Developers should test real backlog items instead of toy prompts. A small syntax edit reveals little about agent loops, cache reuse, or recovery from a wrong approach.

Useful candidates include a bug with a reliable reproduction, a feature spanning several files, or a migration with a defined test suite.

One run is not enough. Repository state, tool latency, and nondeterministic model behavior can alter the route through a task.

Three or four paired tasks provide a more credible initial sample. Larger teams should group results by task type, rather than report one blended average.

Teams must also define success consistently. A run that produces plausible code but fails tests should not count as a cheaper completion.

Human review time belongs in the operational analysis, even when it does not appear in the token receipt. A confusing patch can consume engineering time after generation ends.

Claude Code’s final report may help reviewers understand longer runs. Anthropic presents clearer closing reports as another potential source of efficiency.

That benefit is plausible but workload-dependent. Teams should measure whether reviewers need fewer follow-up prompts or spend less time reconstructing the agent’s actions.

Independent public evidence remains limited because Opus 5.5 launched only four days before this analysis. Early user reports cannot yet establish a stable industry average.

The defensible conclusion is narrower. Opus 5.5 has lower published rates, and cache-heavy tasks receive a larger structural advantage.

Whether the completed task reduction approaches Anthropic’s estimate depends on turns, output, cache behavior, retries, and migration quality.

How to Measure Your Own Claude Code Task Cost

The `/usage` command turns the pricing claim into a repeatable test using your actual sessions.

Run /usage when a coherent task finishes. /cost provides the same view within Claude Code.

The session block reports input, output, cached input, and an estimated cost based on list rates. Subscription users should treat that estimate as a work indicator.

It is not an additional subscription invoice. Plan limits and token-billed API usage represent different billing arrangements.

Start by recording the model and effort setting. Without those details, two session receipts may look comparable while representing different operating modes.

Next, record the task definition and acceptance test. A clear completion condition prevents one run from stopping earlier than another.

Then inspect the cache share. A long session should usually reuse a large portion of its input.

A low cache share can point to pauses, model changes, effort changes, or prompt modifications. It can also reflect a naturally short or fragmented task.

Compare total input with the largest observed context. If total input is many times larger, the session probably used numerous turns.

That difference is not automatically waste. Multi-step engineering work naturally needs several requests, especially when tests reveal new information.

Still, repeated inspection of the same files can identify an avoidable loop. Review the transcript around those repetitions to find missing instructions or validation tools.

Output deserves a separate check because it includes internal thinking. High output on a small mechanical change may indicate excessive effort or repeated reasoning.

For a paired test, reset the repository to the same starting state. Run the task with Opus 5 and then Opus 5.5, alternating order across later tests.

Record turns, fresh input, cache reads, output, elapsed time, test results, and required human corrections. These fields explain the result better than one total.

Use the calculator only after collecting those measurements. Entering real token quantities produces a useful rate comparison.

Keep its efficiency control at zero for the first calculation. That shows how the published rate changes affect the same token workload.

Then calculate the observed token difference from paired runs. This second view combines pricing with the model’s actual behavior.

Do not assume all future tasks will match that sample. Separate debugging, feature work, code review, repository search, and unattended agent runs.

Effort should also be tested by task category. Medium may fit scoped daily work, while difficult failures may justify high effort.

Low effort can suit deterministic edits, but only when verification makes errors cheap to detect. A failed low-effort attempt weakens any apparent saving.

Teams using gateways need an additional check. The gateway must preserve prompt-caching and usage fields, or internal reports can misstate session economics.

Anthropic’s gateway documentation describes centralized usage tracking, budgets, and request attribution. It also warns that outdated gateways can block newer features.

Larger organizations can use usage reports to aggregate results by developer and model. Per-user totals alone remain insufficient without task outcomes.

A useful internal metric pairs completed tasks with token consumption. Another tracks retries that occur after failed tests or reviewer rejection.

Teams can store short experiment notes beside engineering decisions. A searchable technical knowledge base can preserve prompts, settings, outcomes, and migration findings.

That record helps distinguish model changes from process changes. It also prevents every team from repeating the same benchmark without shared methodology.

The goal is not to optimize every session to the smallest receipt. It is to identify settings that deliver accepted code with predictable cost and review effort.

Three Signals Will Show Whether the Savings Hold

The next test is whether lower rates translate into stable, verified output across ordinary engineering work.

The first signal is paired /usage data from real projects. Repeated reductions across debugging, feature work, and review would strengthen the cost-per-task argument.

Those comparisons should publish token categories and success criteria. A headline percentage without cache share, effort, and retries cannot explain what changed.

The second signal is cache stability during long sessions. Teams should watch whether Opus 5.5 maintains high cache reuse across tool calls, compaction, and model transitions.

Consistently low cache shares would weaken the expected advantage. They would suggest that workflow design or infrastructure prevents users from reaching the favorable cache rate.

The third signal is retry frequency after migration. Opus 5.5 changes effort defaults, thinking behavior, and several tool interfaces.

Fewer failed loops would support Anthropic’s claim that the model completes work more efficiently. More integration errors could temporarily erase the published savings.

Developers should also resist reducing the comparison to Opus 5 alone. Smaller Claude models can remain more suitable for search, log reading, and inexpensive summaries.

The relevant decision is workload placement. Opus 5.5 may serve as the main model for supervised coding while smaller models handle bounded retrieval.

Hard, unattended tasks can justify a more capable model if it avoids multiple failures. The cheapest successful route can begin with a higher-cost model.

Anthropic’s calculator improves this discussion by exposing the variables behind the estimate. It does not settle the result for a specific team.

The Claude Opus 5.5 cost per task is lower under equal token usage, especially when cache reads dominate input. The exact reduction remains an empirical question.

Pick one genuine backlog item, define its passing test, and run it once on each model. Compare /usage, turns, output, cache share, and review corrections.

Repeat that process across several task types before changing a team-wide default. If Opus 5.5 finishes with fewer retries, the rate reduction compounds.

If it consumes more reasoning or disrupts an integration, the headline saving will narrow. The next month of paired production measurements will matter more than any single calculator preset.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page