top of page

OpenRouter Agent Token Usage Is 5x Human Traffic, but the Lead Is Mostly Cached Context

4 days ago
15 min read

OpenRouter agent token usage reached 7.3 trillion tokens in August, more than five times the platform’s 1.4 trillion attributed to humans. Yet more than 85% of those agent tokens reportedly came from cached prompts, not newly processed instructions or generated answers.

That distinction complicates the claim that AI now uses more AI than people do. Agents clearly create far more model traffic per task. However, the chart measures tokens routed through one platform, not global AI adoption, spending, productive work, or economic value.

Daniel Newman, CEO of Futurum Group, amplified the numbers in a September 30 X post. He predicted the ratio would rise from five times to 10 times, then continue higher. The five-times figure comes from observed OpenRouter traffic. The 10-times figure remains a forecast without a stated timeline or supporting model.

The real contest is therefore not agents versus humans. It is raw token volume versus useful work. The distinction matters for developers managing agent loops, enterprises assessing returns, and infrastructure providers planning memory capacity.

OpenRouter Agent Token Usage Crossed a Clear Threshold

The August data shows a major change in model traffic, but only within OpenRouter’s measured environment.

OpenRouter operates a gateway that routes requests across models and inference providers. Its position gives the company visibility into diverse application traffic, including direct conversations, coding tools, and autonomous workflows.

The reported chart presents seven-day average token usage through August 10, 2026. Agentic traffic reached approximately 7.3 trillion tokens, compared with 1.4 trillion for human traffic.

The resulting ratio is about 5.2 to one. It supports the narrower statement that agents generated five times more token traffic than humans on OpenRouter during that measurement period.

It does not establish the same ratio across ChatGPT, Claude, Gemini, private cloud deployments, or locally hosted models. Those systems represent substantial traffic that OpenRouter cannot observe.

OpenRouter also reported that agentic usage had grown about fourteenfold since February 6. Human token usage increased about 2.8 times during the same period. Mixed traffic, which combines human and agent-like behavior, reportedly grew by approximately 4.7 times.

February 6 was the last observed date when human traffic remained above agent traffic in the dataset. That makes the August result more meaningful than a single daily spike. Agents had maintained and widened their lead for roughly six months.

OpenRouter classified each API key as agentic, human, or mixed. The system reportedly used seven weighted signals, including tool-call rates, turn counts, and timing gaps between responses.

This approach is more informative than relying only on application names. A generic API client can run an autonomous loop, while an application marketed as an agent can remain under direct human control.

However, behavioral classification introduces uncertainty. API keys can serve multiple products, teams, or use cases. A workload can also shift between human-led and automated behavior without changing credentials.

The mixed category acknowledges this ambiguity, but it does not remove it. Reassigning part of that traffic would change the ratio between agents and people.

There is another important semantic issue. Agents are not independent customers in the usual economic sense. Humans and organizations deploy them, define objectives, fund their requests, and decide whether their outputs have value.

The agents are better understood as automated intermediaries. One human request can initiate planning, retrieval, tool execution, validation, correction, and repeated model calls.

That multiplication effect is the core event. AI demand is no longer determined only by how many people open a chat window. It increasingly depends on how much machine activity each human request triggers.

OpenRouter’s earlier 100-trillion-token study identified the same structural shift. It described agentic inference as an extended sequence involving planning, tools, revisions, and repeated model interaction.

The study found that programming had become a leading source of prompt growth. Programming prompts also averaged several times the length of general-purpose prompts by late 2025.

Those patterns help explain why agents crossed the threshold. A person might submit one coding request. An agent can repeatedly inspect files, call tools, review test results, and resend its working context.

The August chart captures that amplification at platform scale. It does not prove that autonomous software has replaced human demand. It shows that human demand increasingly arrives through systems that create many downstream requests.

One Human Instruction Can Trigger Thousands of Model Operations

Agents consume more tokens because they turn a single request into a continuing computational process.

A conventional chatbot interaction usually follows a visible rhythm. A person writes a prompt, receives an answer, and decides whether to continue. Each new turn depends on another human action.

An agent can continue without that pause. It interprets a goal, creates steps, chooses tools, evaluates results, and decides whether another attempt is needed.

Consider a software migration. A developer might ask an agent to move an application from one cloud service to another.

The first model call can examine the request and form a plan. Later calls might inspect configuration files, search documentation, edit code, run tests, diagnose failures, and revise the implementation.

Each call often includes more than the newest instruction. It can carry system rules, tool definitions, repository details, previous messages, command output, and the agent’s prior decisions.

This accumulated material is the context window, meaning the text and structured data available to the model during one request. As the task continues, that context can become much larger than the original human prompt.

Tool use further increases traffic. A model might generate a database query, inspect the result, and then call the model again to decide what follows.

Parallel agents can multiply the process again. One coordinator might delegate research, coding, testing, and review to separate workers. Every worker maintains its own instructions and task history.

This explains why token volume can grow much faster than the number of users. The basic unit has shifted from a conversation turn to a workflow step.

OpenRouter’s data also reflects growth in reasoning and programming workloads. These tasks naturally support longer interactions because they involve intermediate state, external tools, and repeated validation.

Academic evidence suggests that the amplification can become extreme. One 2026 study of agentic coding found that agent tasks consumed roughly 1,000 times more tokens than code chat in its experimental setting.

The agent-cost research also found wide variation between repeated runs. Token use for the same task differed by as much as thirtyfold in the researchers’ tests.

Crucially, more tokens did not consistently produce better results. Performance often peaked at an intermediate level before additional consumption stopped delivering corresponding gains.

That finding strengthens the main tension behind OpenRouter agent token usage. A rising count can represent productive automation, unnecessary repetition, or a mixture of both.

Agent builders therefore face pressure to measure work completed, not just activity generated. Useful indicators include accepted code changes, resolved support cases, successful transactions, and tasks completed without human repair.

A token is only a unit of text processing. It carries no built-in measure of accuracy, difficulty, novelty, or business value.

Two workflows can consume the same number of tokens while producing very different outcomes. One might resolve a difficult engineering issue. The other might repeat a failed plan until a limit stops it.

The design of the surrounding system determines which outcome is more likely. Clear completion tests help an agent recognize success. Bounded retry policies prevent a failing task from running indefinitely.

Good tools also reduce the need for lengthy textual reasoning. A structured API can return a precise result that would otherwise require repeated browsing and interpretation.

Memory architecture matters for the same reason. An agent does not need every old detail during every step. It needs the subset that remains relevant to the current decision.

This is where a searchable personal knowledge base can support human-led workflows. Retrieved evidence can replace indiscriminate history when the system selects context carefully.

The pressure extends beyond developers. Enterprise buyers now need to evaluate how products control loops, retrieve context, and report consumption.

A product can look efficient during a short demonstration but behave differently across long-running workloads. Production tasks encounter missing permissions, ambiguous goals, changing data, and unexpected tool responses.

Those conditions generate retries. They also expose whether the agent has reliable stopping rules or merely continues producing plausible next steps.

Human adoption still matters, but it no longer predicts inference demand by itself. The more useful formula includes users, delegated tasks, model calls per task, and tokens per call.

That formula makes a ten-times ratio plausible in principle. It does not make Newman’s forecast inevitable. The ratio will also depend on optimization, model behavior, application design, and traffic outside OpenRouter.

Cached Prompts Explain Most of the Token Explosion

The largest part of the agent lead appears to be repeated context, not entirely new reasoning or output.

More than 85% of agent tokens in the a16z presentation reportedly came from cached prompts. Cached prompts are previously processed input segments that a provider can reuse when a later request begins with matching content.

An agent often resends stable material. That material can include system instructions, tool descriptions, project files, policies, and earlier conversation history.

Processing the same prefix from scratch on every call would waste computation. Prompt caching lets the provider reuse intermediate results associated with that prefix.

OpenRouter’s cache telemetry separates cached tokens from newly processed prompt tokens. It can also identify tokens written into a cache for future reuse.

This means 7.3 trillion tokens should not be interpreted as 7.3 trillion units of fresh model work. A large share represents information the system has encountered before.

The distinction affects cost. Providers generally charge less for a cache read than for processing new input because much of the earlier computation has already occurred.

However, cached does not mean free. The system must identify the cache entry, retrieve its data, and make the associated model state available during inference.

The relevant model state is often called a key-value cache, or KV cache. It stores attention information derived from earlier tokens so the model can continue without recalculating every preceding position.

Longer prompts create larger KV caches. More concurrent agent sessions create more of them. Long-lived workflows can also require repeated access to substantial stored context.

That shifts the infrastructure bottleneck. Compute remains important, but memory capacity, memory bandwidth, and data movement become increasingly significant.

This is why the cached share does not make the OpenRouter data meaningless. It changes what the data means.

The chart is weaker as evidence of fresh reasoning demand. It is stronger as evidence that agent systems repeatedly carry large histories across many model calls.

The difference resembles repeatedly consulting the same large project binder. Reading familiar pages again requires less preparation than analyzing new pages, but the binder must remain available.

Caching can also make inefficient agent design financially tolerable. A workflow might resend an enormous system prompt because the discounted cache read hides part of the cost.

That design still consumes capacity. It can increase latency, complicate routing, and create dependence on stable prompt prefixes.

Cache hits are not guaranteed across every configuration. Changing an early part of the prompt can invalidate later cached material. Routing requests between providers can also affect reuse.

Dynamic data presents another challenge. If timestamps, retrieved documents, tool results, or user-specific details appear near the beginning, they can reduce prefix stability.

Agent developers therefore need deliberate context structure. Stable instructions belong near the beginning, while changing information should appear after reusable sections when provider behavior allows it.

Session affinity can matter as well. Requests that remain associated with a compatible provider are more likely to reuse prior context than requests routed unpredictably.

The 85% figure also deserves careful attribution. It appeared in the a16z framing of OpenRouter data, according to the source report. OpenRouter’s original analysis has also been described as showing a lower share under another calculation.

Different denominators can produce different percentages. One method can aggregate all token volume, while another averages the cached share across requests.

A few enormous workflows can dominate total volume without representing the typical request. Conversely, an average per-request percentage can understate the influence of the largest workloads.

Both measurements can be accurate while answering different questions. The public chart does not provide enough methodological detail to reconcile every reported cache percentage independently.

The safest conclusion is therefore directional. Cached context forms the clear majority of agent token traffic, while the precise share depends on how the platform aggregates requests.

This also challenges the phrase “AI is using AI.” The agents are not necessarily performing trillions of independent acts of reasoning. Much of their traffic involves restoring the context required to continue delegated work.

That behavior can still generate real value. A coding agent needs repository state and prior decisions to avoid starting from zero after every tool call.

The efficiency question concerns selection. Does the agent reload the smallest useful context, or does it repeatedly send everything because that approach is easier to implement?

As token volume rises, that difference becomes a material engineering choice. Efficient context management can reduce memory demand without weakening task performance.

Five Times the Tokens Does Not Mean Five Times the ROI

Token volume measures utilization, while return on investment depends on successful outcomes and total operating cost.

Newman argued that companies focus too heavily on human adoption when assessing AI returns. His broader point has merit because one user can now initiate far more inference than a chat-based adoption metric reveals.

Monthly active users can understate infrastructure demand. Seat counts can also miss automated work that runs continuously after employees leave their desks.

Yet replacing user counts with token counts creates another incomplete measure. Tokens reveal activity, but they do not show whether that activity created revenue, reduced labor, improved quality, or increased risk.

The five-times ratio also compares two traffic categories rather than two economic actors. Agent requests remain downstream of human or organizational decisions.

An enterprise does not earn a return because an agent consumed more tokens. It earns a return when the agent completes valuable work at an acceptable cost and risk level.

A useful assessment begins with task success. Teams should ask whether the workflow completed the intended action and whether a human accepted the result.

The next question concerns intervention. An agent that finishes without oversight has a different operating profile from one that requires repeated corrections.

Latency also matters. A workflow that eventually succeeds can still fail commercially if customers must wait too long or if infrastructure queues grow under peak load.

Then comes total cost. Token charges are only one component. Tool calls, search services, databases, sandboxes, observability, security reviews, and human remediation can add substantial expense.

Risk-adjusted outcomes matter as well. An agent that modifies production systems needs stronger controls than an agent summarizing public documents.

The OpenRouter chart cannot answer any of those questions. It was designed to describe traffic, not enterprise returns.

The same limitation applies to Newman’s 10-times prediction. Extrapolating the ratio assumes agent traffic continues growing faster than human traffic without a comparable efficiency correction.

That assumption can fail for several reasons. Applications can compress context, use smaller models for routine steps, replace repeated reasoning with deterministic software, and stop loops earlier.

Model improvements can reduce retries. Better tool interfaces can also return cleaner information, lowering the number of calls needed to finish a task.

Economic pressure will encourage those changes. Businesses have an incentive to remove calls that do not improve outcomes, even when cached reads are relatively inexpensive.

The a16z discussion about loop convergence highlights the same problem. An agent can keep producing additional work after most of the available value has already appeared.

A loop without an external completion test can confuse continued activity with progress. It might repeatedly edit a document, rerun a failing command, or refine an already acceptable answer.

This behavior is especially difficult to detect when each individual call appears reasonable. Waste emerges across the full trajectory rather than within one response.

Observability must therefore operate at the task level. Developers need traces that connect each model call to tool use, state changes, errors, and eventual outcomes.

Budgets should also reflect task value. A high-stakes investigation can justify more iterations than a routine formatting request.

Escalation rules create another boundary. When the agent encounters repeated failure or uncertain permissions, handing control to a person can be cheaper and safer.

There is also a selection effect in OpenRouter’s traffic. The platform serves developers who deliberately use a model gateway, which can produce a more technical workload mix than consumer applications.

Programming and agent frameworks can therefore occupy a larger share of OpenRouter than they do across the entire AI market.

OpenRouter’s scale still makes the trend important. Its data covers substantial real-world traffic across many models and providers. The results should simply remain attached to that scope.

The platform’s relationships introduce another consideration. Andreessen Horowitz has invested in OpenRouter, giving a16z an interest in the growth of token-routing infrastructure.

That does not invalidate the figures. It makes transparent methodology and independent replication more important, especially when the data supports broad claims about the AI economy.

The strongest interpretation avoids both extremes. The chart is neither proof that autonomous machines have become AI’s primary customers nor an empty artifact of caching.

It is evidence that agent architectures amplify inference demand. It also shows that this amplification currently relies heavily on transporting and retrieving old context.

For buyers, the central question is not whether an agent uses many tokens. It is whether each additional round increases the probability of a successful, valuable outcome.

Memory Providers and Agent Platforms Face the Immediate Pressure

The traffic shift rewards systems that manage context efficiently and pressures products that treat token consumption as a proxy for progress.

Model providers face a more complicated workload than ordinary request-response chat. Agent sessions can remain active longer, call tools repeatedly, and maintain growing histories.

That workload places pressure on schedulers and routing systems. Providers must balance cache locality, model availability, latency, and reliability across changing demand.

Gateway platforms such as OpenRouter gain strategic importance because applications increasingly use several models. A workflow can route planning, coding, validation, and summarization to different endpoints.

Dynamic routing can reduce costs or improve performance, but it can also interfere with caching. A request sent to a different provider might lose access to a previously established cache.

Providers that expose clear cache metrics will have an advantage with sophisticated buyers. Teams need to know how many tokens were new, cached, generated, or written into storage.

Memory manufacturers also face demand from larger contexts and more concurrent sessions. High-bandwidth memory feeds model accelerators, while conventional memory and storage support surrounding systems.

However, the chart does not quantify future memory purchases. It shows token traffic, not an exact mapping between each token and new hardware capacity.

Hardware demand depends on model architecture, data types, batching, cache eviction, compression, and the number of simultaneous sessions. Software improvements can change each relationship.

Agent platforms face pressure from another direction. Customers will increasingly compare useful work per token, not just access to capable models.

Products that hide consumption behind broad usage claims can struggle when enterprise finance teams demand task-level economics. Buyers will want repeatable evidence across their own workloads.

Coding agents offer an early test because their outputs can be evaluated with builds, tests, reviews, and deployment outcomes. These signals create measurable stopping conditions.

Other domains remain harder. Research, strategy, and writing often lack a single objective test. Agents can continue refining outputs without a clear moment of completion.

That uncertainty makes human review more important, even when agents perform most intermediate work. It also makes token efficiency harder to compare between vendors.

Infrastructure vendors can respond with context compression, prefix caching, stateful sessions, and retrieval systems. Each approach attempts to avoid resending unnecessary history.

Application developers can split durable memory from working context. Durable memory stores potentially useful information, while working context contains only what the current step requires.

This separation reduces repeated text and limits irrelevant information. It can also improve model accuracy by keeping distractions outside the active prompt.

Deterministic software should handle deterministic work. An agent does not need to reason through a calculation, schema validation, or access check when reliable code can perform it directly.

The most efficient systems will probably combine models with conventional software. Models handle ambiguity and planning, while code enforces rules and performs stable operations.

This hybrid design weakens the assumption that agent traffic must rise without limit. Better systems can complete more tasks while reducing tokens per task.

At the same time, falling unit costs can increase total demand. When each workflow becomes cheaper, developers can deploy agents in more places and run them more often.

The resulting rebound can preserve infrastructure growth even as individual tasks become more efficient. That possibility supports the direction of Newman’s forecast, though not its precise ratio.

Human traffic can also grow. Better consumer products, new interfaces, and broader enterprise adoption could raise direct usage alongside agent traffic.

The future ratio depends on which curve grows faster. It is not solely a function of improving agents.

Competition among OpenAI, Anthropic, Google, open-weight model developers, and specialized inference providers will shape that curve. Their caching rules and tool capabilities differ.

Model choice can also change during a workflow. A small model might classify a request, while a larger model handles a difficult decision.

That architecture reduces the relevance of a single aggregate token count. A token processed by a compact model has a different resource profile from one processed by a frontier model.

Enterprises will need normalized measures that combine model choice, token type, latency, energy, and task success. No widely accepted standard currently captures the entire picture.

Until one develops, OpenRouter agent token usage remains a useful leading indicator. It reveals the shape of demand before financial reporting can explain its value.

Three Signals Will Test the 10x Prediction

The next phase should be judged by classification stability, fresh-token growth, and completed work per unit of inference.

The first signal is whether OpenRouter’s agent-to-human ratio continues rising after August. A sustained increase would support the idea that autonomous workflows are expanding faster than direct interaction.

The comparison should use the same classification method and measurement window. Method changes could create apparent growth without an equivalent change in behavior.

Mixed traffic deserves special attention. If it grows faster than both categories, the boundary between human and agent use will become less reliable.

The second signal is the composition of agent tokens. A rising cached share would indicate that context repetition, not fresh model processing, remains the main growth engine.

A falling cached share paired with rising total volume would tell a different story. It would suggest that agents are performing more new inference rather than mainly replaying established context.

Public reporting should distinguish new input, cache reads, cache writes, reasoning tokens, and output. Combining them into one number conceals major differences in cost and infrastructure demand.

The third signal is task efficiency. Agent vendors and enterprise users should report completed work per model call, tokens per successful task, and human interventions per completion.

If those measures improve while total agent traffic rises, the growth case becomes stronger. It would indicate that adoption and successful automation are expanding together.

If token use rises while success remains flat, the chart would increasingly describe operational waste. A higher ratio would then weaken, rather than strengthen, the ROI argument.

Independent datasets would also improve confidence. Traffic from direct model APIs, cloud platforms, and private deployments could show whether OpenRouter reflects the wider market.

For now, the evidence supports a narrower conclusion than the viral claim. Agents generated more than five times the human token traffic observed on OpenRouter during August 2026.

Most of that traffic appears to involve cached context. That context still requires memory, routing, and careful management, but it should not be confused with an equal amount of new reasoning.

The move toward 10 times is a forecast, not an established outcome. Its importance will depend on what the additional tokens accomplish.

Developers should examine whether each loop has a clear purpose, budget, and stopping condition. Enterprise buyers should demand task-level evidence before treating consumption as adoption.

The useful question is no longer whether agents will generate more traffic than people. OpenRouter’s data suggests they already do. The question is whether the next trillion tokens will complete more work, or simply reread more of the past.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page