Barclays Traces Every $100 in AI Revenue, and Cloud Providers Take Nearly $40
Barclays has mapped every $100 earned by AI model companies and found that cloud providers collect roughly $35 to $40 through inference costs. That estimate places AWS, Microsoft Azure, and Google Cloud near the center of the industry's emerging profit structure.
The finding comes from an AI unit economics report dated August 28, 2026. A Chinese report published on August 30 summarized the analysis and its two hypothetical AI laboratory models.
The headline looks like a straightforward win for cloud infrastructure. However, Barclays also expects the model companies' margins to improve as inference becomes more efficient and training consumes a smaller revenue share.
That creates the central conflict. Cloud providers currently receive a large portion of model revenue, yet their share can decline while the AI market keeps expanding.
This is not simply a contest between software companies and infrastructure vendors. It is a question of which layer keeps the economic advantage as AI moves from expensive training into continuous production use.
What Barclays Found Inside Every $100 of AI Revenue
The cloud bill remains one of the largest claims on AI model revenue, even after substantial gains in inference efficiency.
Barclays built two hypothetical frontier laboratories to show how business models and accounting policies reshape reported margins. These are analytical models, not disclosed financial statements from named companies.
Lab A receives about 70% of its revenue from APIs and 30% from subscriptions. Lab B reverses that mix, receiving 80% from subscriptions and 20% from APIs.
For each $100 earned by Lab A, the associated cloud provider receives about $35 in revenue. Barclays estimates that approximately $11.80 becomes operating profit, representing a margin near 34%.
Lab B produces a different result. Its cloud partner receives roughly $41 from every $100, including a strategic revenue-sharing arrangement.
The cloud provider retains about $19.10 as operating profit in that scenario. The resulting margin reaches approximately 47%.
Those examples underpin the widely repeated conclusion that cloud providers capture about $35 to $40 per $100 of AI laboratory revenue. They also retain roughly $10 to $20 as operating profit.
The exact outcome depends on more than computing consumption. Product mix, partner agreements, training allocation, and revenue recognition all change the apparent distribution.
Lab A recognizes indirect API revenue on a gross basis. Lab B either records comparable revenue on a net basis or excludes business operated by its partner.
Gross recognition records more of the transaction as revenue before associated expenses. Net recognition records only the company's retained portion.
Barclays compares this accounting divide with the historical difference between Uber and Lyft. Similar underlying transactions can produce sharply different financial statements when companies report them differently.
That warning matters because private AI companies provide limited standardized financial data. Comparing headline revenue or gross margin without adjusting for recognition policies can create a false picture.
The two laboratories also carry different adjusted gross margins. Barclays assigns Lab A a margin near 55% and Lab B about 38%, a difference of 17 percentage points.
The API-heavy laboratory performs better partly because direct API customers pay according to token consumption. That structure connects revenue more closely with the computational work performed.
Subscription products create another risk. A fixed fee must cover users with very different consumption patterns, including intensive coding sessions and long-running agent tasks.
Usage limits can protect the provider, but they also affect retention. Raising limits may attract customers while shifting more inference cost back onto the model company.
Barclays estimates that subscription products such as Claude Code and Codex generate inference margins around 70%. Direct API inference margins reportedly exceed 80%.
These figures describe inference contribution, not the complete profitability of an AI company. They exclude or separately allocate major expenses such as training, research, sales, and administration.
The $100 framework is therefore a map, not an audited industry average. Its value comes from showing where money moves and which assumptions change the outcome.
Why Inference Economics Are Improving Now
AI companies are gaining margin because each useful response requires less infrastructure, while commercial demand supports continued monetization.
Inference means running a trained model to answer a request, generate code, or complete an agent task. It becomes a recurring expense each time customers use an AI product.
Training follows a different pattern. It concentrates vast computing costs before a model reaches customers, although post-training and continual improvement blur that boundary.
Barclays estimates that paid inference margins rose from low double-digit percentages in 2025 to between 50% and 65%, or higher, during 2026. Adjusted gross margins improved by 30 to 50 percentage points.
Several technical changes support that shift. Quantization reduces the numerical precision used by a model, cutting memory and computing requirements while trying to preserve output quality.
Speculative decoding lets a smaller system propose tokens that a larger model can verify in groups. Successful predictions reduce the time and hardware needed for generation.
Better routing also sends simpler requests to smaller models. Caching prevents repeated computation, while batching allows hardware to process several requests together.
Hardware utilization matters just as much as model design. An expensive accelerator sitting idle still creates depreciation, financing, energy, and facility expenses.
The largest operators can spread variable workloads across regions, customers, and products. This scale gives cloud providers an important advantage over smaller AI laboratories managing dedicated clusters.
However, model developers also benefit from efficiency. Every improvement in tokens per unit of computing can expand their margin or support more usage at the same cost.
Google reported that it reduced Gemini serving unit costs by 78% during 2025 through model optimization, utilization improvements, and infrastructure efficiency. Its cloud results illustrate how vertically integrated operators can improve several layers together.
Microsoft described a similar focus during fiscal 2026. The company reported a 40% improvement in inference throughput for its most-used Copilot models during one quarter.
Its later infrastructure update said Copilot workload throughput had risen fourfold since the year's start. Microsoft attributed those gains to optimization across silicon, systems, and software.
Those improvements do not automatically lower customer bills. Providers can retain efficiency gains, use them to improve output quality, or support more demanding reasoning workloads.
Reasoning models often generate more intermediate tokens before producing an answer. Agents can call models repeatedly while searching, writing code, checking results, and revising their work.
Consequently, lower unit costs can produce higher total consumption. This rebound effect makes inference efficiency valuable to both AI companies and cloud providers.
Enterprise demand adds another layer. Barclays analyst Ross Sandler linked improving margins to enterprise customers and agentic workflows becoming important purchases.
That assessment remains a forecast rather than a universal customer verdict. Enterprises still vary widely in adoption, deployment scale, security requirements, and measurable returns.
Yet production agents can consume more than model tokens. They require databases, storage, networking, identity management, monitoring, and durable task state.
Those supporting services strengthen the cloud position. A laboratory may optimize model inference while its customers purchase a wider infrastructure bundle from the same cloud platform.
This dynamic explains why the cloud can collect nearly $40 while AI laboratories still improve their margins. Efficiency enlarges the available pool before competition determines who keeps it.
Cloud Providers Win the Current Round, but Fund the Entire Contest
AWS, Azure, and Google Cloud capture recurring AI revenue because they absorb the capital burden before model companies can serve customers.
The Barclays model focuses on operating flows, but the infrastructure layer begins with enormous upfront commitments. Data centers require land, power, cooling, networking, processors, and long construction schedules.
Microsoft reported capital expenditures of $37.5 billion in its fiscal second quarter of 2026. Roughly two-thirds involved shorter-lived assets, primarily GPUs and CPUs.
The company also said customer demand continued to exceed available supply. Its cloud performance showed Azure revenue growth alongside pressure from continued AI infrastructure investment.
Microsoft's Intelligent Cloud revenue increased 29% during that quarter. Operating income rose 28%, although the segment's gross-margin percentage declined year over year.
That combination captures the cloud tradeoff. AI demand produces rapid revenue growth, but satisfying it requires additional equipment that can depreciate faster than buildings or land.
Alphabet planned between $175 billion and $185 billion of capital spending during 2026. It said the money would support Google DeepMind, consumer products, and Cloud customer demand.
Google Cloud ended 2025 with an annual revenue run rate above $70 billion. Alphabet also reported a $240 billion cloud backlog and strong demand for AI products.
Amazon faces the same investment cycle. Its shareholder materials say much of the expected 2026 AWS capital spending will support capacity monetized during 2027 and 2028.
Amazon also said it had customer commitments covering a substantial portion of planned capacity. Its shareholder filing presents that demand visibility as support for continued spending.
These companies can finance infrastructure using cash generated by established businesses. Search advertising supports Alphabet, enterprise software supports Microsoft, and commerce helps fund Amazon.
Standalone model companies lack the same diversified base. They often rely on equity financing, strategic cloud partnerships, advance commitments, or other structured arrangements.
That dependence gives cloud providers economic leverage beyond their per-token margins. They control capacity allocation when demand exceeds supply and influence which models reach enterprise customers.
Marketplaces and managed model platforms deepen that role. A company can purchase model access through an existing cloud relationship instead of contracting directly with every laboratory.
The cloud provider then controls billing, identity, governance, regional availability, and technical integration. The model becomes one component inside a broader enterprise account.
Still, the nearly $40 figure should not be confused with pure economic rent. Cloud vendors must pay for chips, energy, maintenance, networking, staff, financing, and hardware replacement.
Barclays estimates operating profits of $10 to $20 from each $100 of laboratory revenue. The remaining cloud receipts fund those substantial infrastructure costs.
The model also says partner revenue sharing inflates the apparent cloud margin in Lab B. Excluding that arrangement, the cloud provider's profit per token matches Lab A.
That distinction changes the interpretation. The stronger Lab B result does not necessarily show superior infrastructure economics.
Instead, it partly reflects a negotiated transfer between strategic partners. Barclays expects that particular revenue-sharing mechanism to phase out after 2028.
The current cloud advantage therefore rests on three foundations: scarce capacity, large capital commitments, and contracts that connect model developers with infrastructure providers.
All three can change. Capacity can become more abundant, specialized operators can enter, and large laboratories can finance dedicated infrastructure outside traditional cloud contracts.
What the $100 Model Does Not Prove
Barclays offers a useful economic framework, but its projections depend on private-company estimates, hypothetical laboratories, and accounting choices.
The underlying Barclays research was summarized publicly, but the complete analyst model is not broadly available as an audited data set. Readers should treat its figures as estimates.
The two laboratory profiles are deliberately simplified. Real companies combine consumer subscriptions, enterprise agreements, direct APIs, cloud marketplaces, licensing, and strategic partnerships.
They also negotiate different hardware rates and capacity terms. A company with a multiyear commitment can face economics unlike those of a smaller developer buying capacity on demand.
Model architecture creates another variation. Some systems use sparse mixture-of-experts designs, while others activate more parameters for each request.
Workload shape also matters. Short classification tasks, long coding sessions, image generation, and autonomous research agents impose very different costs.
Even the meaning of a token differs across model families and modalities. Comparing nominal token prices cannot reveal the cost of completing the same useful task.
Training allocation produces another source of uncertainty. A laboratory can expense, capitalize, or distribute development costs differently across reporting periods and product lines.
Barclays estimates that training spending represented about 48% of AI laboratory revenue in 2026. It projects that share falling to 35% in 2027 and 30% in 2028.
The direction is plausible because revenue can grow faster than training programs. However, frontier competition may also encourage larger and more frequent training runs.
Independent research found that the cost of the most compute-intensive training runs grew about 2.4 times annually from 2016. The training cost study projected billion-dollar frontier runs if that historical trend continued.
That work predates Barclays' 2026 model and measures a different question. Still, it shows why declining training costs as a revenue percentage are not guaranteed.
Revenue can also become harder to compare as indirect APIs grow. One laboratory may record the full customer payment, while another records only its retained share.
This can make one company appear larger or more profitable without reflecting a better underlying product. Barclays specifically warns that future comparisons require careful accounting adjustments.
Strategic relationships create an additional distortion. A cloud company can invest in a model developer while also receiving infrastructure spending or revenue sharing from that developer.
These arrangements are commercially legitimate, but they complicate any clean division between supplier revenue, investment returns, and independent customer demand.
Cloud-level financial reporting cannot fully resolve the problem. AWS reports a distinct segment, but Microsoft does not disclose Azure revenue as a separate dollar figure.
Google Cloud includes infrastructure, platform services, applications, and other enterprise offerings. None of these segments isolate profits from frontier-model inference.
The reported 34% to 47% margins in the Barclays examples therefore cannot be checked directly against public cloud accounts. They remain modeled allocations.
Competition can pressure both sides. Model companies may reduce prices to defend usage, while cloud vendors may discount capacity to win large commitments.
Open models can also shift demand toward infrastructure without preserving the original developer's revenue. Enterprises can host those models through a cloud platform or specialized provider.
At the same time, proprietary laboratories can route workloads across multiple vendors. They can use custom accelerators, dedicated clusters, or alternative clouds to improve bargaining power.
The headline should therefore remain narrow. Barclays estimates that cloud providers currently collect nearly $40 from every $100 in model-company revenue under representative structures.
It does not prove that every laboratory pays the same percentage. It also does not establish that cloud companies will retain that share indefinitely.
The Three Signals That Will Decide Who Keeps the Next $100
The next phase will depend on training intensity, dedicated infrastructure, and the gap between cloud investment and realized AI revenue.
The first signal is training spending as a percentage of laboratory revenue. Barclays expects that ratio to decline sharply through 2028.
If the ratio reaches the projected 35% during 2027, model companies will have stronger evidence that revenue is outrunning frontier development costs.
That outcome would strengthen the case for improving laboratory profitability. It would also reduce the cloud's total receipts relative to each dollar of model revenue.
If training intensity remains near 2026 levels, the opposite conclusion follows. The industry would still depend heavily on costly development cycles despite better inference margins.
The second signal is the arrival of dedicated capacity outside the three dominant cloud platforms. Barclays expects the traditional cloud share to remain broadly stable for two years.
After 2028, previously contracted or financed infrastructure should become a larger computing source for AI laboratories. That forecast deserves close examination well before the facilities open.
Announcements alone provide weak evidence. Investors and customers should watch whether projects secure power, chips, financing, network connections, and operational teams.
They should also track whether laboratories move production workloads onto that capacity. Training clusters do not automatically become reliable, geographically distributed inference systems.
If dedicated facilities handle substantial production traffic, the near-$40 cloud capture rate should decline. Model companies would internalize more infrastructure responsibility and potential margin.
However, they would also absorb utilization risk and capital exposure. Owning capacity only improves economics when workloads keep expensive hardware consistently productive.
The third signal is the relationship between hyperscaler capital spending and cloud profitability. Microsoft, Alphabet, and Amazon are committing vast sums before all capacity becomes revenue-producing.
Strong cloud growth accompanied by stable or improving margins would support Barclays' argument. It would show that infrastructure vendors can convert AI demand into durable operating profit.
Slower growth or persistent margin compression would weaken it. Such results would suggest that competition, depreciation, or energy costs consume more value than the $100 model assumes.
Enterprise adoption provides the connecting thread. Cloud companies need more than a few frontier laboratories to justify infrastructure built on this scale.
Microsoft said demand exceeded available supply during fiscal 2026. Alphabet reported that AI customers used more products than non-AI cloud customers, expanding revenue beyond raw computing.
Those claims will become more credible if customers renew commitments and move pilots into production. Usage must persist after early experimentation budgets expire.
For developers and enterprise buyers, this economic contest affects product design. A workflow with repeated model calls can create a larger infrastructure burden than its interface suggests.
Teams should evaluate the full task cost, including retries, retrieval, storage, monitoring, and human review. The cheapest token is not always the cheapest completed result.
Knowledge workers face a similar issue. Agent subscriptions can look simple while hiding variable infrastructure consumption behind usage limits and changing service policies.
Keeping an internal record of results, decisions, and source material can reduce unnecessary repeated requests. A searchable AI knowledge base also helps teams reuse completed work.
Barclays projects global AI laboratory revenue rising from $137 billion in 2026 to $690 billion in 2028. That forecast remains aggressive and should not be treated as settled fact.
Even a lower outcome could produce substantial cloud demand. The decisive question is not whether the market grows, but which layer converts that growth into retained profit.
Watch the next financial disclosures for standardized revenue recognition, inference costs, training allocation, and contractual commitments. Better disclosure can test assumptions that hypothetical laboratories cannot settle.
Also watch whether AI developers continue celebrating margin improvements while signing larger infrastructure agreements. That combination would show that efficiency gains are stimulating more consumption.
The near-$40 figure is best understood as a snapshot of bargaining power in 2026. Cloud providers finance scarce infrastructure, while model companies control products and customer relationships.
Which side keeps more of the next $100 will depend on execution, not a permanent rule. Follow those three signals before treating today's distribution as tomorrow's industry structure.



