top of page

Anthropic Cuts Claude Sonnet 5.5 Cache Pricing, Matching OpenAI at $0.10

5 hours ago
10 min read

Anthropic has cut Claude Sonnet 5.5 cache pricing by 50%, lowering cache reads from $0.20 to $0.10 per million tokens. The company says the change makes Sonnet 5.5 about 20% cheaper for most agentic work, where applications repeatedly reuse long instructions and context.

That qualification matters. Anthropic did not reduce the model’s standard input or output rates. It changed one component that becomes important when an agent repeatedly sends the same system prompt, tool definitions, repository context, or reference material.

The move also brings Sonnet 5.5’s cache-read rate in line with OpenAI’s GPT-6.1 Sol. Both models now list the same standard input, cached-input, and output prices. That turns the competition away from headline token rates and toward a harder question: which model completes a reliable task with fewer requests, fewer generated tokens, and less rework?

Claude Sonnet 5.5 Cache Pricing Falls by Half

The immediate change is narrow, but it targets one of the largest recurring costs in long-running agent workflows.

Anthropic announced the reduction on October 7, 2026, alongside the release of Claude Haiku 5.5. Its model announcement says Sonnet 5.5 cache reads now cost $0.10 per million tokens instead of $0.20, effective that day.

A cache read occurs when an application reuses previously stored prompt content. Instead of processing that content again at the full input rate, the provider retrieves the matching prefix and charges the lower cache-read rate.

The feature matters most when requests contain a large, stable prefix. A coding agent might repeatedly send repository instructions, tool descriptions, coding standards, and selected source files. A research system might reuse a lengthy policy manual or document collection across many questions.

Customer-support agents can follow the same pattern. They often carry a stable system prompt, product catalog, escalation rules, and tool schemas into every turn. Only the customer’s newest message and the agent’s latest working state change.

Under Sonnet 5.5’s current pricing, Anthropic lists standard input at $2 per million tokens and output at $10 per million. A cache hit now costs 5% of the standard input rate, down from 10%.

Cache writes remain more expensive than ordinary input because the reusable content must first be stored. Anthropic lists five-minute cache writes at $2.50 per million tokens and one-hour writes at $4 per million. The lower read price pays off only when an application reuses the stored prefix enough times.

Consider a simplified workload that sends 100,000 stable tokens across 100 requests. Without caching, those repeated tokens represent 10 million standard input tokens. With an effective cache, the first request writes the prefix, while later requests retrieve most of it at the cache-read rate.

The new rate halves the read portion of that example. It does not reduce the cost of new input, cache creation, model output, external tools, retries, or failed tasks.

Anthropic’s estimated 20% reduction therefore describes what the company calls “most agentic work.” It is not a universal discount on every Sonnet 5.5 request. Short prompts, low cache-hit rates, or output-heavy workloads will see a smaller change.

The price cut follows Sonnet 5.5’s September 28 launch. In its original Sonnet 5.5 release, Anthropic priced cache reads at $0.20 and said the model usually required fewer tokens than Sonnet 5.

That launch claim already positioned efficiency at the task level, not merely the token level. Anthropic said Sonnet 5.5 could cost up to 30% less per task than its predecessor and generate output more than 30% faster.

The cache adjustment adds another efficiency lever less than two weeks later. It also sharpens Anthropic’s message that persistent agents, rather than isolated chatbot responses, are becoming the unit that matters.

Why Long-Running Agents Make Cache Hits More Valuable

Agents repeatedly consume the same context, so cache economics can matter more than a modest change in ordinary input pricing.

A conventional chatbot request may include a brief instruction and one user message. An agentic application usually carries much more machinery into each model call.

That machinery can include an operating policy, dozens of tool definitions, task history, retrieved records, software documentation, and a plan assembled over earlier steps. Many components stay unchanged while the agent searches, edits files, calls services, and verifies its result.

Prompt caching stores a reusable prefix from that request. Later calls can retrieve the matching content at a reduced rate instead of billing every token as fresh input.

Caching does not mean the model remembers information permanently. It is a billing and processing mechanism for matching prompt content within a defined cache lifetime. If the prefix changes, expires, or fails to meet the provider’s caching rules, the application must write it again.

That distinction creates three practical variables.

First, developers need a high cache-hit rate. Stable instructions and tools should appear before frequently changing messages. Unnecessary edits to the cached prefix can invalidate reuse.

Second, the application needs enough repeat calls to recover the cache-write premium. A prefix used once offers no read savings. A prefix used hundreds of times can make the read rate a major cost component.

Third, output generation still matters. Sonnet 5.5 output costs $10 per million tokens, 100 times its new cache-read rate. An agent that generates long explanations, repeats itself, or enters retry loops can erase the gains from cheaper cached context.

This is why the percentage reduction varies by application. Anthropic’s 20% figure implies a workload where cache reads account for a substantial but not dominant share of the original bill.

A simple cost model makes that relationship clearer. Suppose cache reads previously represented 40% of a workload’s model spending. Cutting that component in half lowers the total by 20%. If cache reads represented only 10%, the same price change would reduce total spending by 5%.

Neither example predicts a particular customer’s bill. They show what must be true for Anthropic’s average claim to hold.

The largest beneficiaries are likely to be systems with long, reusable prefixes and many sequential calls. Coding agents fit that pattern because repository instructions and tool definitions often persist throughout a task.

Document analysis can also benefit. A team might place a large contract, technical specification, or research packet into a cached prefix, then ask a series of focused questions.

Knowledge-work agents have similar needs. They may repeatedly consult stable company policies, meeting records, or project documentation while producing a report. Teams building these systems still need disciplined context selection, much like maintaining a searchable engineering knowledge base.

Caching weakly organized context does not make it useful. It only makes repeated transmission cheaper. Poor retrieval can still fill the prompt with irrelevant material and cause the model to spend output tokens resolving ambiguity.

Anthropic’s own prompt caching documentation recommends placing static content near the beginning of a prompt and dynamic content later. That structure increases the chance that later requests will match a stored prefix.

Developers should also monitor cache creation and read counters separately. A low read-to-write ratio can reveal short cache lifetimes, unstable prefixes, routing changes, or requests that never reuse their stored context.

The new Claude Sonnet 5.5 cache pricing makes those engineering decisions more valuable. It does not remove the need to measure them.

Anthropic and OpenAI Now Share the Same Headline Rates

The cut neutralizes a visible OpenAI price advantage and forces buyers to compare the cost of completing whole tasks.

OpenAI introduced GPT-6.1 Sol with the same $2 input and $10 output rates that Anthropic charges for Sonnet 5.5. However, GPT-6.1 Sol launched with cached input at $0.10 per million tokens.

Before Anthropic’s reduction, an application comparing only those three rate-card fields saw a clear difference. Sonnet 5.5’s cached input cost twice as much.

That difference has now disappeared. Sonnet 5.5 and GPT-6.1 Sol both list:

  • Standard input at $2 per million tokens

  • Cached input at $0.10 per million tokens

  • Cache writes at $2.50 per million tokens for the basic cache duration

  • Output at $10 per million tokens

OpenAI says GPT-6.1 Sol’s cached input costs 5% of its standard input rate. Its model documentation also lists a 1.05-million-token context window and support for multiple reasoning-effort settings.

The matching rates make a simple price comparison less useful. Two agents can use models with identical token prices and still produce very different bills.

One model might solve a coding issue in eight calls. Another might need 15 calls, generate more reasoning tokens, invoke extra tools, or repeat an unsuccessful patch. The second system can cost more despite an identical rate card.

Reliability changes the calculation again. A cheap first attempt offers little value if a person must identify an error, restore damaged work, and rerun the task. The economically relevant measure is the cost of an accepted result.

Latency also has a cost. An agent that completes work with fewer sequential steps can free compute capacity and shorten user waiting time. That can matter more than a small difference in token spending for interactive products.

Anthropic is trying to frame Sonnet 5.5 around that task-level measure. The company reports a 70.6% result on Terminal-Bench 4.0, which evaluates multi-step professional tasks in a command-line environment.

Anthropic also says Sonnet 5.5 runs more than 30% faster than Sonnet 5 and can use fewer tokens for the same work. These are company-reported results, and production performance depends on the agent harness, prompts, tools, and task distribution.

OpenAI makes its own performance and efficiency claims for GPT-6.1 Sol. Its launch announcement describes the model as approaching GPT-6 Astra’s capabilities at lower standard token rates.

Neither company’s benchmark suite provides a universal winner. Anthropic’s tests use particular effort settings and agent configurations. OpenAI’s evaluations use its own research environment and acknowledge that API behavior can differ.

For enterprise buyers, the practical comparison now requires a representative evaluation set. Teams need to measure successful completion, human acceptance, tool calls, cached tokens, uncached input, output, latency, and retries.

They should also test the same operational constraints. Giving one model extra tools, a different context package, or a more generous reasoning setting makes the cost comparison unreliable.

The primary contest is no longer Sonnet’s cache price against OpenAI’s. It is Anthropic’s claim of efficient task completion against the reality of each customer’s production workload.

The 20% Savings Claim Has Important Limits

Anthropic’s estimate is plausible for cache-heavy agents, but it should not be treated as an automatic reduction in total operating costs.

The first limitation is cache eligibility. Applications only receive the lower price when requests produce actual cache hits. Similar content is not necessarily identical content.

Moving a tool definition, changing a timestamp, reordering retrieved documents, or modifying an early system instruction can break the reusable prefix. Small architectural choices can therefore decide whether the advertised rate applies.

The second limitation is cache lifetime. Anthropic supports different storage durations with different write prices. A system with long pauses between requests may need repeated cache writes, reducing its net savings.

The third limitation is output cost. Cache reads are now inexpensive, but generated tokens remain much costlier. Long reasoning traces, verbose responses, and repeated corrections can dominate the final bill.

The fourth limitation is orchestration overhead. Agents often call search systems, browsers, databases, code execution environments, or third-party APIs. Anthropic’s model-price reduction does not lower those charges.

The fifth limitation is human review. A model that produces inexpensive but unreliable work can increase labor costs. This risk is especially important for software changes, regulated workflows, and actions that affect customer data.

Even Anthropic’s original Sonnet 5.5 announcement cautions that benchmarks capture only one aspect of model capability. The company says Opus 5.5 remains stronger for complex, open-ended tasks requiring sustained judgment.

That creates a routing tradeoff. Developers can assign routine, well-scoped work to Sonnet 5.5 while reserving harder decisions for a larger model. However, incorrect routing can cause Sonnet to fail repeatedly before the system escalates.

A poorly designed router can waste more money than the cache discount saves. Teams should measure the full route, including failed attempts and escalation calls.

The comparison with GPT-6.1 Sol has another complication. Matching headline prices do not mean the providers implement caching identically.

OpenAI’s caching guide says newer models support explicit or implicit breakpoints, depending on the model. It also describes minimum prompt lengths and reporting rules that affect which tokens qualify.

Anthropic uses its own cache controls, durations, breakpoints, and usage reporting. Migrating an agent between providers may require prompt restructuring before its cache-hit rate becomes comparable.

Cloud distribution can complicate the picture further. Sonnet 5.5 is available through Anthropic and partner platforms, but pricing, regional availability, and feature timing can vary by provider.

The announced $0.10 rate should therefore be verified on the specific endpoint used in production. A company running through a cloud marketplace should not assume that every regional service adopted the change simultaneously.

There is also a measurement question behind the 20% claim. Anthropic has not published a detailed distribution showing cache consumption across customer workloads. “Most agentic work” groups together coding, research, browser use, customer support, and other patterns with different token profiles.

The safest interpretation is straightforward. Anthropic reduced one verified unit price by half. Its broader percentage represents the company’s modeled or observed workload mix, not a guaranteed customer outcome.

Teams can validate the claim by comparing several weeks of usage. The useful measures include cache-hit tokens per request, cache-write frequency, uncached input, output, successful tasks, and total spending per accepted task.

A drop in cache spending without a similar drop in task cost would signal another bottleneck. That bottleneck might be output length, failed attempts, external tools, or human correction time.

What Developers and AI Buyers Should Watch Next

The next phase will test whether lower cache rates improve production economics or simply become the new baseline for competing agent platforms.

The first signal is Anthropic’s updated billing data. Developers should confirm that their Sonnet 5.5 requests receive the new cache-read rate and that partner platforms expose the same change.

This strengthens Anthropic’s case if customers see lower spending per completed task without changing their applications. It weakens the broader 20% claim if most workloads show only a small reduction.

The second signal is competitive pricing. OpenAI’s matching GPT-6.1 Sol rate already appears to have removed room for Anthropic to charge twice as much for cached context.

Google and other model providers now face the same pressure. Buyers increasingly expect cached input to cost a small fraction of ordinary input, especially for agents designed around persistent context.

A further price response would show that cheap context reuse has become a standard competitive feature. No response would suggest providers still differentiate through model quality, infrastructure, or distinct caching systems.

The third signal is independent task-level testing. Token prices reveal how providers calculate a bill, but they do not show how much work a model needs to finish a job.

Useful evaluations should report cost per accepted result, not only benchmark accuracy or price per million tokens. They should also disclose effort settings, tool access, retry policies, cache-hit rates, and human-review criteria.

For developers, the immediate action is to inspect actual usage rather than multiply token prices by a theoretical prompt size. Segment cache reads, cache writes, uncached input, and output for representative tasks.

Then connect those figures to outcomes. Track whether the agent completed the task, whether a person accepted it, and whether the system needed escalation or repair.

Cache optimization should begin with the largest stable prefixes. Keep system instructions and tool definitions consistent, place dynamic content later, and avoid injecting changing metadata before reusable material.

Teams should also test cache expiration behavior. A five-minute cache can work well for dense agent loops but poorly for workflows with long approval pauses. The more expensive one-hour option may lower total costs when it prevents repeated writes.

The final purchasing decision should compare at least two models on the same internal task set. Matching Claude Sonnet 5.5 cache pricing and GPT-6.1 Sol pricing makes that experiment easier to interpret, but it does not decide the winner.

Anthropic’s reduction is meaningful because agent workloads reuse context far more often than ordinary chat sessions. It also confirms that model vendors now see cached context as a competitive battleground.

The open question is whether lower cache prices produce cheaper successful work. Measure your cache-hit rate, retries, output volume, and accepted results for the next month. If total cost per completed task falls near Anthropic’s estimate, the cut changed your economics. If it does not, the expensive part of your agent lies somewhere else.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page