David Cahn Asks: Can AI Answer the $3 Trillion Question Before the Payoff Slips?
- Ethan Carter

- 6 hours ago
- 12 min read
Sequoia partner David Cahn has raised his AI infrastructure estimate to $1.5 trillion, forcing investors to confront a much larger revenue problem. Can AI answer the $3 trillion question before the expected payoff slips beyond 2028?
That number represents Cahn’s estimate of the revenue needed to justify the industry’s 2026 infrastructure spending. It covers the economic burden behind chips, data centers, power systems, networking equipment, and operator margins.
The conflict is no longer simply AI optimists against skeptics. It is the industry’s promised revenue against the physical and financial commitments already made by Google, Meta, Microsoft, Amazon, and their partners.
AI companies are growing quickly, but infrastructure commitments have grown on another scale. Falling token costs and cheaper open-weight models make adoption easier while making the revenue target harder to reach.
That reversal puts hyperscalers under pressure. Investors expect their enormous capital programs to produce sharply higher free cash flow during the next several years.
If that cash does not arrive on schedule, the consequences extend beyond model developers. They reach chip suppliers, power companies, credit markets, major stock indexes, and the wider economy.
David Cahn Turns a $200 Billion Warning Into a $3 Trillion Test
Cahn’s new estimate changes the AI debate from a question about product demand into a test of whether revenue can catch infrastructure spending.
Cahn began making this argument in 2023, when Nvidia was reporting annual GPU revenue near $50 billion. He added the operating costs and margins required across the infrastructure chain.
That calculation suggested the industry needed roughly $200 billion in revenue to repay the initial investment. The estimate became known as AI’s "$200 billion question."
The question was not meant as a declaration that AI would fail. Cahn framed it as a challenge for entrepreneurs to build applications that could use the available computing capacity profitably.
Three years of aggressive construction have changed the scale. According to a July 9 infrastructure analysis, Cahn now estimates 2026 AI infrastructure spending at $1.5 trillion.
His model implies that the industry needs $3 trillion in revenue to justify those expenditures. The figure reflects more than the purchase price of accelerators.
Data centers need power, cooling, networking, memory, land, construction labor, and ongoing maintenance. Infrastructure operators also expect a return on their capital.
Cahn argues that the final requirement might be higher. Memory constraints, specialized inference chips, and rising construction costs can increase the revenue needed for each gigawatt of capacity.
A gigawatt measures power capacity, not AI output. It matters because electricity has become one of the clearest limits on how quickly companies can deploy new computing systems.
Apollo has published a separate infrastructure forecast that illustrates the same pressure. It estimated that supporting AI infrastructure through 2028 would require nearly $3 trillion in capital.
Apollo expected around half of that requirement to come from outside financing. Its capital forecast also estimated that private credit could provide more than $800 billion.
The estimates measure different things and should not be treated as identical. Cahn is testing the revenue required to justify spending, while Apollo is estimating infrastructure financing needs.
Both calculations point toward the same underlying problem. AI’s physical buildout is advancing faster than the industry has demonstrated a matching stream of durable customer revenue.
The expansion also has long lead times. Cahn has said an average AI data center takes roughly two years to build.
Projects announced during the initial generative AI surge are therefore arriving after many purchasing decisions were made. Their economic value depends on demand that must persist for years.
Cloud infrastructure offered a slower precedent. Amazon introduced AWS in 2006, and the transition of corporate workloads unfolded over decades.
That pace allowed providers to fund expansion progressively. The current AI cycle asks companies to secure chips, electricity, and construction capacity before demand becomes equally visible.
The $3 trillion question is therefore not asking whether people use AI. Usage has already spread across search, software development, customer service, research, content production, and office work.
It asks whether customers will pay enough, for long enough, to support the infrastructure assembled on their behalf. That is a much stricter test.
Can AI Answer the $3 Trillion Question With Today’s Revenue?
The leading AI companies have built substantial businesses, but their reported revenue remains far below the level implied by Cahn’s calculation.
Anthropic was reportedly approaching $60 billion in annual recurring revenue by July 2026. Annual recurring revenue, or ARR, projects current subscription and contract revenue across a full year.
OpenAI reportedly generated $13 billion during 2025. The company said in November 2025 that it had reached $20 billion in ARR, suggesting continued acceleration beyond recognized annual revenue.
Those figures show real demand, particularly for coding agents and general-purpose assistants. They do not close a multi-trillion-dollar gap.
Revenue from model providers is also only one portion of the answer. Google, Microsoft, Amazon, and Meta can monetize AI through cloud services, advertising, commerce, subscriptions, and internal efficiency.
Google can apply models to search advertising and cloud workloads. Microsoft can place assistants inside business software while selling computing capacity through Azure.
Amazon can use AI in its retail operations and offer model infrastructure through AWS. Meta can improve advertising systems and user engagement without charging directly for every model interaction.
These indirect returns complicate Cahn’s calculation. A dollar of improved advertising performance can help justify infrastructure even when it never appears as model API revenue.
Cost savings matter too. An agent that shortens a software project or automates customer support can create value without becoming a separate product line.
However, indirect benefits are difficult to isolate. Companies might credit AI for improvements that also reflect pricing changes, organizational cuts, market growth, or conventional software upgrades.
The strongest near-term applications remain concentrated. In a 2026 AI outlook, Cahn identified coding and ChatGPT as the two established killer applications.
He expected both categories to approach or exceed tens of billions in annual revenue. He also saw a growing group of startups moving toward significant commercial scale.
Coding agents offer an unusually clear economic case. They can draft code, review changes, test interfaces, investigate failures, and operate development tools.
The buyer can compare those results with engineering time, release speed, and defect rates. That creates a more measurable return than a general promise of improved knowledge work.
Even here, revenue attribution remains challenging. A company might pay for several assistants while engineers continue using existing development tools and cloud services.
The relevant question is not whether the agent produces useful code. It is whether the resulting productivity exceeds subscription, inference, review, security, and integration costs.
Enterprise implementation introduces another delay. Large organizations must connect models to private data, permissions, software systems, and approval processes.
They also need evaluation methods that distinguish a fluent response from a reliable outcome. A model that performs well in a demonstration can still fail within a long production workflow.
Cahn has described enterprises as fatigued by difficult internal AI projects. That fatigue creates an opening for application startups, but it also slows the revenue needed by infrastructure owners.
A focused application can package models around one job and remove integration work. Yet the application provider still faces pressure from falling model costs and rapid feature copying.
The revenue case must therefore operate across several layers. Model providers need contracts, application companies need retention, and hyperscalers need profitable utilization of their data centers.
Higher utilization alone is insufficient if every unit of work becomes cheaper faster than demand expands. Revenue depends on the interaction between usage volume and declining unit cost.
For AI to answer the $3 trillion question, customers must consume dramatically more intelligence as its cost falls. That outcome is plausible, but the required increase is enormous.
Cheaper Tokens Create AI’s Central Revenue Reversal
The same efficiency gains that make AI easier to adopt can weaken the revenue growth needed to repay its infrastructure.
Model companies generally charge for tokens, the small units of text or code processed during generation. Lower token costs reduce the expense of running assistants and agents.
That is good news for developers and enterprise buyers. A workflow that was too expensive last year can become commercially practical after a model or hardware improvement.
However, lower unit costs pressure suppliers. If an agent performs the same task with fewer tokens, the provider receives less usage revenue unless customers run more tasks.
This is the central reversal in Cahn’s argument. Technical progress strengthens the product while raising the volume required to recover infrastructure investments.
OpenAI said GPT-5.4 needed fewer tokens for many tasks than GPT-5.2. The company also positioned it as a stronger model for coding, tool use, and professional workflows.
On one public coding evaluation, GPT-5.4 scored 57.7 percent, compared with 55.6 percent for GPT-5.2. OpenAI’s model evaluation also reported large gains in agentic web search.
OpenAI later claimed its newest model was 54 percent more token-efficient on agentic coding tasks. That claim describes the amount of computation expressed through tokens, not a guaranteed productivity gain.
Independent production results can vary with prompts, tools, task length, and reasoning settings. A benchmark cannot establish the final cost of an enterprise deployment.
Even so, the direction is clear. Models are competing not only on intelligence, but also on how efficiently they convert computation into completed work.
Efficiency can stimulate demand through a familiar economic effect. When a resource becomes cheaper, people often discover more ways to use it.
Developers may move from occasional code completion to agents that monitor tests continuously. Support teams may analyze every customer interaction instead of a sample.
Researchers may run several parallel investigations. Product teams may generate, test, and compare many more prototypes before choosing one.
These scenarios can increase total token consumption despite lower costs per task. The infrastructure thesis requires that expansion to happen across a very large market.
Agentic workflows offer one possible route. An agent does not answer a single prompt and stop. It plans, calls tools, reads files, checks results, and repeats steps.
Research published in 2026 found that agentic coding tasks could consume far more tokens than ordinary code chat. The same study found wide variation between repeated runs.
That variation helps explain why efficiency claims require caution. More tokens did not consistently produce better accuracy, and models struggled to predict their own consumption.
Buyers will increasingly demand completed work rather than raw token volume. That shifts pressure toward outcome-based evaluation, predictable budgets, and auditable execution.
A company organizing its own AI work can benefit from keeping sources, decisions, and model outputs together. A searchable AI knowledge base can reduce repeated retrieval and preserve the context behind results.
Better context can lower waste, but it does not guarantee reliability. Teams still need human review for consequential financial, legal, medical, security, and operational decisions.
The economic question is whether these workflows generate enough additional activity to offset efficiency gains. Suppliers need consumption to expand faster than unit costs decline.
That requirement resembles other computing transitions. Cheaper storage created more data, and cheaper bandwidth produced video streaming and cloud applications.
AI may follow the same pattern. The difference is the amount of capital committed before the new demand curve has fully emerged.
Open-Weight Models Put Hyperscaler Forecasts Under Pressure
Open-weight competition threatens the assumption that rising AI usage will flow primarily through the most expensive American model and cloud platforms.
An open-weight model makes its trained parameters available for users to download or operate. It does not necessarily disclose its training data or full development process.
Organizations can run these models on their own infrastructure, use specialized hosting services, or adapt them for narrower tasks. That flexibility can reduce dependence on one provider.
Apollo chief economist Torsten Slok identified this shift as a risk to hyperscaler forecasts. He noted that Chinese models were gaining share while token prices continued falling.
According to Apollo’s July analysis, Chinese models led their American counterparts in token usage among the 20 most-used models. The underlying charts drew from OpenRouter and other usage datasets.
Such datasets do not represent the entire market. They can overrepresent developers who switch frequently between models and underrepresent private enterprise deployments.
Still, they reveal meaningful competitive behavior. Users often choose a cheaper model when it performs well enough for a routine task.
A software team might reserve a frontier model for architecture decisions while using smaller models for tests, documentation, classification, and formatting.
A customer service operation can route simple requests to a local model. It can escalate ambiguous conversations to a stronger hosted system.
This mixture reduces the amount any single provider can charge. It also makes benchmark leadership less valuable when several models meet the buyer’s required threshold.
Google, Meta, Microsoft, and Amazon remain protected by scale, distribution, and existing customer relationships. They can bundle AI with cloud services, productivity software, advertising systems, and developer platforms.
They also have the balance sheets to continue building during a price war. Smaller infrastructure providers lack that cushion.
However, scale creates exposure as well as protection. The largest companies have committed the most capital and carry the highest investor expectations.
Consensus forecasts examined by Apollo expect hyperscaler free cash flow to more than double over the coming years. Free cash flow measures cash remaining after operating expenses and capital spending.
Those forecasts assume today’s investment produces larger and more profitable businesses. They also assume the payoff arrives quickly enough to outrun depreciation and financing costs.
Depreciation spreads an asset’s cost across its expected useful life. It becomes particularly important when newer accelerators make existing hardware economically less attractive.
A data center can remain physically functional while generating weaker returns than planned. Faster chips can lower the value of older systems before their accounting life ends.
Cahn has also warned that delayed artificial general intelligence timelines could leave present capacity exposed. If advanced capabilities arrive later, some hardware may age before demand justifies it.
That risk does not mean installed chips become useless. Existing systems can serve smaller models, fine-tuning jobs, inference, research, and lower-priority workloads.
The issue is the return earned relative to the original forecast. Capacity built for scarce premium intelligence may eventually supply a competitive commodity market.
Slok outlined three possible consequences if the payoff disappoints. First, cash flow and earnings forecasts would fall while capital spending and depreciation continued.
Second, a sell-off in the largest technology companies could spread through chip, energy, data center, and equity markets.
Third, hyperscalers might use more debt when operating cash no longer covers expansion. That would raise leverage and expose lenders to more infrastructure risk.
Apollo’s analysis goes further, arguing that a delayed payoff could contribute to a recession and an S&P 500 correction. This is a risk scenario, not an established forecast.
The claim also needs context. The largest hyperscalers have diverse businesses, strong cash generation, and the ability to change spending plans.
They can delay projects, renegotiate supply commitments, shift workloads, or redirect capacity toward internal products. Their investment programs are not completely inflexible.
There is also a bullish reading of open competition. Cheaper models can bring AI into more organizations, increasing demand for cloud storage, networking, security, and deployment tools.
Hyperscalers could earn money from hosting competing models even when their own models lose share. Infrastructure platforms have historically benefited from broad software ecosystems.
The pressure therefore comes from timing and margins, not a simple collapse in demand. Adoption can keep rising while financial returns arrive later than markets expect.
Three Signals Will Show Whether the AI Payoff Is Arriving
The answer will appear first in utilization, free cash flow, and enterprise behavior, not in another headline benchmark.
The first signal is hyperscaler free cash flow. Google, Meta, Microsoft, and Amazon must show that AI investment supports expanding cash generation through 2028.
Revenue growth alone will not settle the issue. Investors need to see operating cash exceed the demands of data center construction, equipment purchases, and supporting infrastructure.
Companies will offer different explanations because their business models differ. Cloud backlog matters for Microsoft, Amazon, and Google, while advertising returns carry greater weight for Meta.
The useful comparison is the relationship between capital spending and incremental cash flow. A widening gap would strengthen the delayed-payoff case.
A narrowing gap would weaken it, particularly if companies identify repeatable AI revenue rather than relying on general engagement claims.
The second signal is physical data center utilization. Cahn has suggested that warehoused chips would indicate construction delays or a mismatch between equipment deliveries and available facilities.
Hardware needs power, cooling, networking, and completed buildings before it can produce revenue. A delivered accelerator sitting unused remains a capital cost.
Watch for delayed opening dates, power constraints, and changes to equipment deployment schedules. Also watch whether providers report that new capacity sells quickly after launch.
Apollo said in July that on-demand GPU capacity remained effectively sold out. It also reported firm rental demand for some older chips.
That evidence supports the near-term shortage thesis. Persistent shortages would weaken the idea that the industry has already overbuilt.
The picture changes if rental rates fall while completed capacity remains empty. That combination would suggest supply is catching demand before revenue reaches Cahn’s required scale.
The third signal is enterprise spending on completed work. Businesses need to move beyond experiments and renew AI systems because those systems deliver measurable outcomes.
Useful indicators include agent completion rates, customer retention, production workload growth, and documented labor or cycle-time savings. Raw prompt counts provide less information.
Software development will remain an important test. Coding has clear workflows, high labor costs, digital tools, and measurable outputs.
If coding agents become standard production infrastructure, they can support sustained demand across models and clouds. If adoption stalls at supervised assistance, revenue expectations need adjustment.
Customer service provides another test. Models must resolve cases correctly, follow policy, protect data, and know when to escalate.
A lower cost per conversation matters only if quality remains acceptable. Otherwise, savings shift into review, correction, and customer recovery work.
Knowledge work presents a harder measurement problem. Organizations must connect assistants to reliable internal context before outputs become consistently useful.
That requires document access, permissions, source tracking, and institutional memory. A practical knowledge workflow can help teams test whether AI saves time across recurring work.
The next several quarters should reveal whether these examples become durable systems or remain scattered productivity features. Durable systems generate recurring demand and resist easy budget cuts.
Investors should also separate model revenue from infrastructure returns. A successful application does not automatically guarantee strong margins for every data center owner.
Likewise, falling API costs do not automatically prove weak demand. They can unlock workloads that were previously uneconomic.
Can AI answer the $3 trillion question? The industry has demonstrated demand, rapid product improvement, and several valuable applications.
It has not yet demonstrated revenue on the scale implied by its infrastructure commitments. The decisive evidence will come from cash, utilization, and repeat enterprise purchases.
Cahn’s calculation is best treated as a stress test rather than a countdown to failure. It identifies the amount of economic value that the current buildout must eventually support.
For developers and enterprise buyers, the practical response is not to stop adopting AI. It is to measure outcomes carefully and preserve the evidence behind each decision.
Track completed work, review costs, model substitutions, and real usage growth. Those records will show whether efficiency is creating more value or merely lowering the bill.
The $3 trillion answer will not arrive in one earnings call. It will emerge as thousands of organizations decide which AI workflows deserve to become permanent.


