Barclays Says AI Labs Send Nearly 40% of Revenue to Cloud Giants
Barclays says AI model companies send roughly $35 to $40 of every $100 in revenue to AWS, Microsoft Azure, and Google Cloud. That estimate turns the AI boom into a more complicated story about who captures its economic value.
The analysis was reported on August 30 by several outlets, including a 36Kr newsflash. It describes inference, the computing performed whenever a deployed model processes a request, as a major transfer of revenue from AI labs to cloud operators.
Barclays reportedly estimates that cloud providers retain $10 to $20 in operating profit from every $100 earned by an AI lab. That implies operating margins between 35% and 45% on the related infrastructure revenue.
The surprising part is not that AI companies have large computing bills. The reversal is that improving model economics have not removed the cloud companies from the value chain. AI labs and hyperscalers are becoming profitable on the same requests, but infrastructure owners still collect their share first.
That creates the central tension behind the numbers. OpenAI, Anthropic, and other model developers want to turn intelligence into a high-margin software product. AWS, Azure, and Google Cloud supply much of the computing capacity required to deliver it.
The $40 Claim Changes How AI Revenue Should Be Read
An AI lab’s revenue can rise quickly while a large and recurring share immediately leaves as infrastructure spending.
The Barclays estimate begins with a hypothetical $100 of AI-lab revenue. According to reports describing the research, approximately $35 to $40 goes toward inference capacity supplied by the three largest cloud platforms.
That payment covers the hardware and systems involved in serving model responses. These systems include accelerators, networking equipment, storage, power, cooling, and the software used to schedule workloads.
The estimate does not mean every AI company pays the same percentage. A laboratory’s costs depend on its model architecture, customer mix, contract terms, utilization rates, hardware, and ability to operate its own infrastructure.
A coding agent that works through a large repository can consume far more tokens than a short consumer chatbot exchange. An enterprise deployment with strict latency requirements can also carry different costs from a background summarization job.
The figure is better understood as an industry model, not an audited statement for one named company. Barclays has reportedly described the analysis as an “AI Labs & Hyperscaler Unit Economics Primer,” but the full institutional report is not publicly available.
That access gap matters. Readers cannot independently inspect every assumption behind the $35 to $40 estimate, including workload volume, hardware depreciation, contracted discounts, and model efficiency.
The public accounts broadly agree on the headline figures. A separate Barclays summary also attributes the $35 to $40 infrastructure share and $10 to $20 operating profit to the bank’s research.
However, repeated reports can still originate from the same underlying document. They should not be treated as separate confirmations of the model’s calculations.
The estimate nevertheless offers a useful framework for understanding AI economics. Revenue reported by a model provider is not equivalent to value retained after inference.
Traditional software can serve an additional user at a relatively small incremental cost. Generative AI must perform substantial computation whenever a user asks for an answer, generates an image, or delegates a task to an agent.
Model developers can reduce that cost through quantization, caching, batching, routing, and smaller specialized models. They can also negotiate cheaper capacity or deploy custom systems.
Yet greater efficiency does not automatically reduce total infrastructure spending. Lower costs can encourage users to make more requests, while agents can generate long sequences of model calls without continuous human input.
That rebound effect changes the significance of falling token costs. The cost of one unit of intelligence can decline while the total number of purchased units grows much faster.
The Barclays model captures this difference between unit efficiency and aggregate demand. AI labs may become better at serving each request while cloud providers continue selling more total computation.
That is why the $40 claim creates real tension. It suggests that the model layer is improving without escaping its reliance on the infrastructure layer.
Why AI Inference Margins Rose So Quickly
The reported margin improvement reflects cheaper model serving, better utilization, and stronger revenue, not the disappearance of computing costs.
Barclays reportedly estimates that margins on paid inference rose from the low double digits in 2025 to approximately 50% to 65%, or higher, in 2026. The reported year-over-year improvement in adjusted gross margin reaches 30 to 50 percentage points.
Paid inference means model usage connected to revenue, such as API calls, subscriptions, or enterprise products. Gross margin measures revenue remaining after the direct cost of providing that usage.
Several mechanisms can move that margin rapidly. New accelerators process more work, while better software keeps those accelerators occupied for a larger share of the day.
Model providers also use caching to avoid recomputing repeated information. Batching combines compatible requests, spreading infrastructure costs across more output.
Quantization reduces the numerical precision used to represent a model’s parameters. That can lower memory and computing requirements, although aggressive quantization can affect output quality.
Distillation trains a smaller model to reproduce selected behavior from a larger one. The resulting model can be cheaper to operate for tasks that do not require maximum capability.
Routing adds another efficiency layer. A service can direct simple requests to smaller models and reserve expensive frontier systems for difficult problems.
Revenue mix matters just as much as technical efficiency. Enterprise contracts, high-value coding workloads, and usage-based products can generate more revenue from a given amount of infrastructure.
The Barclays figures indicate that model companies have made meaningful progress on both sides of the calculation. They appear to be earning more revenue from paid requests while reducing the computing cost attached to each useful response.
However, a gross margin of 50% to 65% does not equal company-level profitability. AI laboratories still pay for training, research, employees, data, safety systems, distribution, and general operations.
Training expenditure is particularly important because the reported model focuses on inference economics. Building the next model generation can require a separate pool of computing resources long before that model generates revenue.
Sales incentives can also complicate the picture. Free usage, bundled access, promotional credits, and minimum-spend contracts may affect the relationship between recognized revenue and actual demand.
The estimated improvement therefore supports a narrower conclusion. Paid model serving is reportedly becoming a viable high-margin activity, even after significant cloud costs.
That is a meaningful change from the idea that every additional AI query necessarily deepens a provider’s losses. It does not establish that every major laboratory has reached sustainable net profitability.
Higher margins can also attract stronger competition. Model providers may cut API rates, include more usage in subscriptions, or spend additional computing resources to improve answer quality.
Barclays reportedly expects inference margins to decline gradually as frontier-model competition intensifies and more computing supply becomes available. That prediction follows normal market logic.
If several providers can deliver comparable intelligence, customers gain bargaining power. Some efficiency savings then pass through to buyers instead of remaining with the laboratory.
The same pressure can appear through product design. Providers may allow longer contexts, more agent steps, and additional reasoning within existing subscriptions.
Those additions improve the product but consume more computation. They can turn hardware efficiency into better service before it becomes lasting profit.
This is the mechanism behind the reported reversal. AI laboratories have improved their margins, yet competition can convert part of that gain into lower prices and heavier workloads.
Cloud Giants Get Paid No Matter Which Model Wins
AWS, Azure, and Google Cloud occupy a favorable position because model competition can increase infrastructure demand across the entire market.
An AI laboratory must persuade customers that its models offer better performance, reliability, or integration. A cloud provider can benefit when several laboratories compete and all require more computation.
This arrangement resembles an infrastructure toll, although the cloud companies also face substantial costs and investment risks. They must purchase equipment and build data centers before future demand becomes certain.
Recent financial disclosures show why investors take the cloud position seriously. Amazon reported that AWS generated $42.2 billion in second-quarter 2026 sales, up 37% from the prior year.
AWS produced $16.6 billion in quarterly operating income, compared with $10.2 billion one year earlier. Those figures imply a segment operating margin of approximately 39%.
Amazon also said its AWS AI business had exceeded a $25 billion annual revenue run rate. The company described that business as growing at triple-digit percentages year over year in its quarterly results.
Those results do not isolate inference revenue from other AWS services. They also do not prove the specific Barclays allocation for AI laboratories.
They do show that a major cloud platform can expand rapidly while producing substantial operating income. AWS accounted for most of Amazon’s segment operating income during the quarter.
Microsoft offers a different view because it does not publish Azure as a separate reporting segment. Its Intelligent Cloud division includes Azure alongside server and enterprise products.
Microsoft said Azure and other cloud services revenue grew 43% in the quarter ending June 30, 2026. Annual Azure revenue surpassed $100 billion and grew 41%, according to its earnings call.
The Intelligent Cloud segment reported $39.3 billion in quarterly revenue and $16 billion in operating income. Its operating margin remained near 41%.
Microsoft also reported that customer demand continued to exceed available cloud capacity. The company added 31 data centers during the quarter and 88 during the fiscal year.
Those figures support the demand side of the Barclays thesis. Capacity remains valuable because customers are ready to use new infrastructure soon after it becomes available.
They also reveal the cost side. Microsoft’s cloud gross margin declined as the company expanded AI infrastructure ahead of demand and supported growing product usage.
Microsoft Cloud’s gross margin fell to 65% in the latest quarter, down from 68% one year earlier. That business includes more than Azure, so the figure cannot represent pure AI infrastructure profitability.
Google Cloud presents a third structure. Google sells cloud infrastructure and applications while also developing Gemini models and custom Tensor Processing Units.
Alphabet reported Google Cloud revenue of $13.6 billion in the second quarter of 2025, up 32%. Operating income reached $2.8 billion, giving the segment a 20.7% operating margin in its cloud results.
The three companies therefore combine several economic roles. They sell raw infrastructure, managed databases, model platforms, enterprise applications, and access to first-party or third-party models.
That breadth increases their opportunity to capture value. A cloud provider can earn revenue from the accelerator, storage, networking, database, security layer, and managed AI service behind one application.
The model developer still holds important leverage. The strongest models attract users, influence application design, and can negotiate large capacity agreements across several suppliers.
OpenAI and Anthropic have expanded their infrastructure relationships beyond a single platform. Greater supplier diversity can reduce operational dependence and improve negotiating power.
Cloud providers are also competing with one another. AWS develops Trainium and Inferentia, Google offers TPUs, and Microsoft designs its own accelerators while continuing to deploy external hardware.
Custom chips can lower costs and differentiate each platform. They can also make workloads more closely tied to a provider’s software and operational environment.
The primary contest is therefore not AWS against Azure or Google Cloud. It is AI laboratories seeking software-like margins against infrastructure owners collecting a recurring share of model revenue.
Competition among the three clouds supports that contest, but it does not replace it. Even a laboratory using several providers still pays for the underlying computing work.
What the Barclays Estimate Does Not Prove
The reported figures describe a plausible industry structure, but they do not provide audited economics for each AI laboratory or cloud platform.
The first uncertainty is the underlying data. Barclays is a major financial institution, yet the full report and its methodology are not available on an open public page.
Public summaries provide the output of the model without exposing every input. That makes it difficult to test the estimate against different token volumes, utilization assumptions, and capacity agreements.
The second uncertainty is accounting scope. AWS reports segment revenue and operating income, but its segment includes far more than AI inference.
Google Cloud also combines infrastructure and Workspace applications. Microsoft’s Intelligent Cloud segment is broader than Azure, while Microsoft Cloud includes several additional products.
Reported cloud operating margins cannot be assigned directly to AI inference. Shared research expenses, corporate costs, and hardware depreciation may also sit in different accounting categories.
A cloud provider could report a strong segment margin while earning less on newly deployed AI capacity. Mature storage, database, and conventional computing services can support the broader result.
The opposite can also occur. A platform may accept lower near-term margins while it fills new capacity and builds a larger customer base.
The third uncertainty concerns internal transfers and strategic relationships. A model company may receive investment from the same cloud provider that sells it computing capacity.
Cloud credits, preferred access, revenue-sharing agreements, and equity investments can blur a simple customer-supplier relationship. The nominal infrastructure bill may not capture the full economics for either side.
The fourth uncertainty is workload variation. A short text request, an image-generation task, and a long-running software agent do not have comparable inference costs.
Models can also spend additional computation before producing an answer. Reasoning systems often trade higher inference spending for better performance on difficult tasks.
As usage shifts toward autonomous agents, the number of model calls per user action can rise sharply. A human may initiate one job while the software performs many searches, evaluations, and revisions.
That makes cost control an application-design problem, not only a model problem. Developers must decide when a higher-quality response justifies additional computation.
Enterprise buyers face a similar issue. A low token price does not guarantee a low total bill when an application runs across thousands of employees and repeated automated processes.
Buyers should monitor cost per completed task rather than cost per token alone. The cheaper model is not always cheaper if it requires more retries, more supervision, or additional downstream work.
This distinction also matters when evaluating productivity claims. A system that spends more on inference can still deliver better economics if it saves enough employee time or prevents expensive mistakes.
Developers building internal systems can preserve the decisions behind those deployments in a searchable technical knowledge base. That record helps teams compare model quality, task completion, and operating cost over time.
The fifth uncertainty is pricing behavior. Falling inference costs do not guarantee expanding margins when providers aggressively reduce prices.
Open models can place additional pressure on proprietary APIs. Enterprises can deploy some open models on rented or privately owned infrastructure, although they then assume more operational responsibility.
Specialized cloud providers can also compete for AI workloads. They may offer attractive accelerator access, flexible contracts, or systems designed around a narrower range of tasks.
Those alternatives weaken the idea that three hyperscalers will permanently capture the same percentage of AI revenue. They do not eliminate the need to finance chips, energy, networking, and data-center capacity.
Barclays reportedly expects actual AI-lab margins to exceed some of its estimates. That remains a modeling judgment rather than a verified disclosure from the laboratories involved.
The estimate should therefore be treated as a map of the value chain. It shows where revenue probably flows, but it cannot settle which company has the best unit economics.
Three Signals Will Test the Cloud Profit Thesis
The next phase will be decided by disclosed margins, infrastructure diversification, and the amount of computation included in AI products.
The first signal is cloud profitability during continued capacity expansion. AWS, Microsoft, and Google must show that rising AI demand can support margins while infrastructure spending remains elevated.
Amazon’s results provide a relatively clean benchmark because AWS publishes segment revenue and operating income. Investors can compare future AWS growth with the segment’s margin and Amazon’s cash expenditure.
Microsoft deserves separate treatment. Azure growth can remain strong even while the broader cloud gross margin falls because Microsoft is building capacity and supporting more AI usage.
The company’s investor metrics reported a 65% Microsoft Cloud gross margin for the latest quarter. Future disclosures will show whether efficiency gains offset the cost of heavier AI adoption.
If cloud margins remain stable while AI-related revenue grows, the Barclays thesis becomes stronger. It would indicate that infrastructure providers can absorb large investments without surrendering their economics.
A sustained margin decline would weaken the thesis. It would suggest that competition, depreciation, energy, and new capacity consume more of the AI revenue than the model assumes.
The second signal is whether leading AI laboratories move meaningful inference workloads beyond AWS, Azure, and Google Cloud. Diversification can include specialized providers, owned facilities, or more extensive custom hardware agreements.
A laboratory does not need to abandon the hyperscalers to gain leverage. It only needs credible alternatives for the next block of capacity.
Long-term contracts will matter more than announcements. The key evidence is where new production workloads actually run, not where a company performs a limited technical trial.
Barclays reportedly expects some infrastructure projects backed by committed customers to take share from the three largest providers. If that happens, the estimated $35 to $40 transfer could become more distributed.
The cloud giants may respond through lower rates, custom chips, financing, or deeper software integration. Each response could preserve workload volume while reducing the margin earned on it.
If major laboratories continue signing their largest production commitments with the same three platforms, the cloud position remains intact. Supplier diversification without workload migration would change little.
The third signal is how much inference providers include in each product. Model companies can preserve margins through usage-based billing, stricter limits, and routing to cheaper systems.
They can instead compete by offering longer contexts, more agent actions, and higher reasoning effort for the same subscription. That strategy can increase customer value while compressing margins.
Enterprise AI purchasing will make this signal easier to observe. Buyers increasingly evaluate successful task completion, latency, security, and predictable spending together.
A product that reduces inference cost without maintaining output quality will not create durable value. A product that improves quality without controlling total consumption can become difficult to scale.
Teams should track model choice, prompts, outputs, approvals, and business outcomes in one repeatable workflow. A structured AI workflow can help reveal whether added model usage actually improves the finished work.
The Barclays estimate ultimately reframes the AI profit debate. Model laboratories are no longer simply burning money on every paid request, according to the reported figures.
Cloud providers are not passive vendors either. They finance and operate the physical systems that make each request possible, then collect a substantial portion of model revenue.
The decisive question is whether AI labs can reduce that infrastructure share faster than competition forces them to spend the savings. Watch the next cloud earnings, major capacity contracts, and changes to product usage limits.
Those signals will show whether nearly $40 of every $100 remains a durable cloud toll or becomes a temporary feature of AI’s infrastructure buildout.



