Epoch AI Cost Report Finds AI Prices Falling Faster Than Moore's Law
Epoch AI says the cost of reaching a fixed AI performance level has fallen 47% per quarter since 2023. The new Epoch AI cost report translates that pace into a 13-fold annual decline, far faster than historical improvements in computing hardware.
That finding does not mean every AI bill falls by the same amount. It measures the cheapest available way to reach a defined score across five benchmarks covering mathematics, science, and games. The distinction matters because newer models can deliver comparable results while consuming fewer resources or using cheaper infrastructure.
The comparison with Moore's Law is still striking. Compute became steadily cheaper over decades, helping make personal computers, smartphones, and cloud services possible. Epoch AI argues that performance-adjusted AI prices are now falling six times faster than the historical rate for compute.
The immediate beneficiaries are developers and organizations that can move workloads between models. The pressure falls on OpenAI, Anthropic, Google, and third-party AI services that need recurring revenue. A premium model can establish a lead, but its economic advantage may disappear within months.
That creates the central tension. Cheaper intelligence supports wider adoption, yet the same decline weakens price-based loyalty and compresses the value of raw model access. The durable business may sit above the model layer, where products retain context, workflow integration, trust, and proprietary data.
What the Epoch AI Cost Report Actually Measured
Epoch AI measured the declining cost of a result, not simply the advertised price of an AI token.
The research organization published its cost analysis on September 22, 2026. Researchers Luke Emberson and David Roodman examined model performance and expense across five benchmarks beginning in 2023.
The benchmarks cover difficult questions in mathematics, graduate-level science, chess, and other games of skill. Epoch AI then estimated the cheapest available route to each attainable performance level at different dates.
This boundary is called the Pareto frontier. It represents the models that provide the best available performance for a given expense. A model outside that frontier costs more without delivering a better result.
The method also accounts for reasoning budgets. A reasoning model can spend more output tokens considering a problem, raising its chance of answering correctly. A simple token price therefore does not reveal the full cost of completing a task.
Epoch AI used evaluation transcripts to estimate how performance changes under different budget limits. The approach follows work by the federal Center for AI Standards and Innovation, or CAISI. It produces an expense-performance curve for each model instead of reducing that model to one advertised rate.
This distinction becomes important when two models use tokens differently. A model with cheaper tokens may generate far more of them before completing a task. Its apparent price advantage can disappear when measured from question to correct answer.
Across the five benchmarks, Epoch AI calculated an average performance-adjusted decline of 47% per quarter. That equals roughly a 13-fold reduction over one year if the trend compounds.
The pace varied by domain. Prices fell about 39% to 43% per quarter on game-based puzzles. Mathematics benchmarks showed faster declines of roughly 50% to 52% per quarter.
One example illustrates the scale without relying on token list prices. Epoch AI compared two OpenAI models released less than 18 months apart. The later model reportedly reached the same GPQA Diamond score at roughly one seven-hundred-and-twenty-fifth of the earlier model's expense.
GPQA Diamond is a multiple-choice benchmark designed around graduate-level questions in physics, chemistry, and biology. It captures one demanding kind of reasoning, but not the full range of useful work.
The headline number is therefore best read as a price index for measured capability. It does not describe training costs, corporate spending, or every customer's invoice. It tracks how cheaply a benchmark result can be purchased from the most cost-effective available model.
That result also aligns with earlier evidence. Stanford's AI Index found a more than 280-fold decline in the cost of GPT-3.5-level performance between November 2022 and October 2024.
The Stanford finding used MMLU, a broad knowledge benchmark, rather than Epoch AI's complete set of expense-performance curves. Different methods produce different rates, but both point toward an unusually steep decline.
The convergence matters more than any single estimate. Independent datasets increasingly show that useful AI capability is becoming cheaper much faster than the hardware underneath it.
Why AI Performance Prices Are Falling So Fast
The decline combines better algorithms, smaller capable models, improved hardware, and aggressive competition among providers.
Moore's Law describes the historical tendency for transistor density to increase over time. The phrase is also associated with falling compute costs, although transistor counts and customer prices are not identical measures.
Epoch AI estimates that compute prices fell at a compounded rate of about 1.51-fold per year between 1940 and 2001. Its estimated AI performance price decline since 2023 is about 13-fold per year.
The difference exists because AI benefits from more than semiconductor progress. Model developers improve architectures, training methods, data selection, quantization, inference software, and hardware utilization simultaneously.
Quantization reduces the numerical precision used to store and run a model. When applied carefully, it lowers memory and compute requirements while preserving much of the model's useful performance.
Mixture-of-experts systems provide another route. They activate only selected portions of a model for each input instead of using every parameter. That design can increase capability without requiring the full network for every request.
Distillation transfers useful behavior from a larger model into a smaller one. The smaller model can then handle routine tasks with lower latency and resource use. Organizations can reserve larger models for problems where their added capability changes the outcome.
Competition magnifies these engineering gains. Providers repeatedly release smaller models designed to approach the performance of earlier flagships. Open-weight models also let independent hosts compete on serving efficiency.
A 2026 research paper called The Price of Progress reached a more conservative conclusion using Epoch AI and Artificial Analysis data. It estimated annual declines of fivefold to tenfold for comparable benchmark performance.
The authors attributed the decline to hardware efficiency, algorithmic gains, and economic competition. After attempting to separate hardware and market effects, they estimated roughly threefold annual progress from algorithmic efficiency.
That estimate is below Epoch AI's 13-fold headline rate. The gap does not make either analysis useless. It shows how strongly the result depends on benchmarks, performance thresholds, model sets, and frontier definitions.
Epoch AI's metric favors the cheapest available model at each capability level. That framework assumes a highly active buyer who regularly reevaluates the market and switches to the current cost leader.
Actual organizations behave differently. They spend time testing models, rewriting prompts, checking security, updating evaluations, and renegotiating contracts. Those switching costs prevent many users from capturing the full theoretical decline.
The study also finds that prices fall fastest immediately after a performance level first appears. Across its five primary benchmarks, near-frontier performance became 66% cheaper per quarter at first.
Two years after a performance level debuted, the estimated quarterly decline slowed to 32%. That still represents a rapid annual reduction, but it is less dramatic than the initial race.
Epoch AI suggests that new capability briefly earns a premium. Competitors then reproduce or approximate it, while the original provider optimizes serving costs. The premium shrinks as the capability becomes common.
This pattern explains why the aggregate decline can outpace Moore's Law. The market does not wait for a new chip generation. It stacks software, model, infrastructure, and competitive improvements on top of continuing hardware gains.
It also explains why the trend can persist while AI companies spend more on data centers. Training the strongest model and serving an established capability are different economic activities.
Frontier laboratories can pour increasing resources into research while customers pay less for yesterday's performance. The industry is simultaneously making the frontier more expensive to create and existing capability cheaper to consume.
Cheaper Intelligence Puts Model Vendors Under Pressure
Rapid price declines turn raw model capability into a depreciating product with a shortening window for premium returns.
OpenAI, Anthropic, Google, and other model developers compete to establish the performance frontier. That position attracts attention, enterprise trials, and developers seeking the strongest available system.
Epoch AI's findings suggest that this lead loses economic value quickly. Across the five benchmarks, the fastest declines appeared around newly achieved performance levels. The premium attached to a leading result then weakened as competitors caught up.
This creates a difficult revenue problem. A laboratory needs substantial investment to train, evaluate, secure, and operate a frontier model. Yet the price customers accept for a fixed capability can fall before that model has enjoyed a long commercial life.
The original news analysis highlighted the resulting challenge for customer loyalty. If a comparable model becomes cheaper within months, buyers gain a reason to keep alternatives available.
Third-party AI services face an even sharper version of this problem. A product that merely resells access to one model can be undercut by another service using a newer model.
Customers may also question a long commitment to any single provider. Today's preferred model might lose its advantage after the next release cycle. A fixed architecture can turn falling upstream costs into migration pressure.
The market response is already visible in product design. Providers increasingly compete through coding environments, agents, file systems, connectors, memory, security controls, and deployment options.
Those features create value that a benchmark does not measure. A model switch becomes harder when the surrounding service understands a company's permissions, terminology, documents, and review process.
Reliability also matters. An enterprise does not benefit from a cheaper answer when that answer triggers more human review. The useful metric is the total cost of completing acceptable work, including failures and oversight.
Latency creates another distinction. Two models may achieve similar benchmark scores while taking different amounts of time to respond. Interactive products often value predictable speed more than a small reduction in inference expense.
Privacy and governance can outweigh price as well. Buyers need to know where data goes, how long providers retain it, and whether administrators can enforce access rules.
These factors give established vendors defenses against pure price competition. Familiar interfaces, existing integrations, security approvals, and accumulated user context all create practical switching costs.
However, those defenses work only when the product provides value beyond model selection. A thin wrapper with generic prompts has little protection when capable substitutes become abundant.
The strongest independent services will likely treat models as replaceable infrastructure. They can route simple tasks to efficient systems and reserve premium capability for difficult work.
That approach requires evaluation rather than loyalty. Developers need representative tests for their own use cases, including quality, latency, failure rates, and end-to-end resource consumption.
It also changes how organizations should view AI procurement. The question is no longer which single model wins every benchmark. The question is which portfolio delivers acceptable outcomes across different tasks.
Knowledge-intensive products have another advantage. A system that organizes a user's trusted information can retain value even when its underlying model changes. The accumulated context belongs to the workflow rather than one model generation.
This is where an AI knowledge base can become more durable than raw chat access. Its value comes from retrieval, organization, permissions, and continuity across repeated work.
Falling model costs can improve margins for such products. They can use stronger models without raising customer-facing costs, or apply AI to more steps inside a workflow.
The same decline can hurt services that compete mainly on access. As artificial intelligence becomes cheaper, differentiation moves toward proprietary context, dependable execution, and measurable business outcomes.
What the Numbers Do Not Prove
Benchmark-adjusted price declines are real evidence, but they are not a universal measure of useful or trustworthy work.
Epoch AI states several limitations directly. Its primary dataset covers about three years, contains incomplete model-benchmark combinations, and includes noisy estimates.
That is a short window for comparing AI with technologies measured over decades. Compute, batteries, sequencing, and electricity also have different products, markets, and definitions of performance.
The comparisons still provide historical scale. However, a benchmark answer is less physically standardized than a unit of electricity or battery capacity.
Benchmark optimization presents another concern. Developers may train models on tasks resembling popular evaluations, intentionally or indirectly. Performance can then improve faster on the test than in unfamiliar real-world settings.
Epoch AI calls this risk benchmaxxing. Randomized benchmark elements can reduce recognition and memorization, but they cannot remove every form of contamination or targeted optimization.
A high score also does not guarantee dependable work. GPQA Diamond measures responses to difficult science questions. It does not measure whether a model can maintain a project, follow company policy, or avoid costly errors.
NIST's evaluation guidance emphasizes that model comparisons depend on task settings, tools, prompts, reasoning levels, sample counts, and scoring methods. Small configuration differences can alter both performance and expense.
CAISI evaluations also show why end-to-end measurement matters. A model with lower token rates can consume more tokens, use more tool calls, or fail more often. Its total task expense can therefore exceed that of an apparently costly rival.
Real users also do not switch models as efficiently as Epoch AI's frontier method assumes. Migration requires testing, integration work, risk review, and staff training.
A company may rationally keep a more expensive provider because it already meets security requirements. Another may value stable outputs because changes would disrupt automated workflows.
Demand can absorb part of the savings. When an AI task becomes cheaper, developers often run it more frequently, expand context, request more candidates, or add verification stages.
This is a version of the Jevons effect. Greater efficiency lowers the cost of each unit, but total consumption can rise enough to maintain or increase overall spending.
Reasoning models reinforce that effect. A more efficient base model can spend extra computation checking an answer or exploring alternatives. Users gain better results without necessarily seeing a proportionate decline in their total usage.
The report also does not show that frontier development has become inexpensive. Training a new leading model requires different resources from serving a known capability.
Research teams pay for experiments, specialized labor, data preparation, evaluation, and infrastructure. Those costs can rise even while inference for established tasks becomes dramatically cheaper.
This distinction resolves an apparent contradiction. Model providers can announce enormous infrastructure programs while the performance-adjusted price of their output continues falling.
The report therefore supports a narrow but important conclusion. Comparable measured capability has become much cheaper at the market frontier.
It does not prove that every workload improves at the same pace. It does not guarantee equivalent safety, reliability, privacy, latency, or integration quality.
Buyers should treat the 47% quarterly rate as a market signal, not a budget forecast. The direction appears well supported. The exact pace remains sensitive to measurement choices and user behavior.
Three Signals That Will Test the AI Cost Decline
The next test is whether benchmark savings survive contact with real workloads, vendor economics, and everyday switching behavior.
The first signal is independently measured end-to-end task cost. More evaluations should compare the total expense required to complete coding, research, customer support, and document workflows successfully.
These tests need stable task sets, controlled tools, and clear quality thresholds. They should include retries, reasoning tokens, tool calls, latency, and human review.
If those measurements fall at a pace near Epoch AI's estimate, the report's commercial implications will strengthen. A much slower decline would show that benchmark savings overstate practical gains.
The second signal is vendor pricing and product packaging. Model providers can respond to falling inference prices by reducing rates, increasing usage allowances, or bundling capability into broader subscriptions.
They can also protect revenue through agent platforms, enterprise controls, storage, and proprietary tools. Those additions would confirm that competition is shifting above the base model.
Watch how often providers replace premium features with standard access. A rapid migration from flagship exclusivity into smaller models would support Epoch AI's finding about fleeting frontier premiums.
The third signal is customer switching. The frontier calculation assumes buyers repeatedly choose the cheapest capable model. Real adoption data will show whether they actually behave that way.
Developers may adopt routing systems that send each task to an appropriate model. Enterprises may demand portability clauses and standardized evaluations before making long commitments.
Alternatively, integrated workflows may keep customers with one provider despite large price differences. That outcome would weaken the claim that falling performance prices automatically destroy vendor loyalty.
Open-weight models will influence all three signals. They give independent hosts more freedom to optimize deployment and challenge closed providers on cost. They also transfer more responsibility for security, maintenance, and evaluation to the operator.
The likely result is not a uniform race to the lowest rate. AI markets will separate into layers.
At the model layer, comparable capability should continue facing price pressure. At the application layer, vendors will compete through context, workflow ownership, trust, and execution quality.
For developers, the practical response is to preserve optionality. Keep model interfaces modular, maintain evaluation sets, and measure successful task completion rather than token consumption alone.
Enterprise buyers should also shorten assumptions about cost stability. A contract that looks efficient today can become uncompetitive quickly, even when the underlying service remains capable.
Knowledge workers should expect AI features to spread into more products. Lower inference costs make summarization, retrieval, drafting, classification, and background processing economical in more situations.
The Epoch AI cost report does not settle how profitable the sector will become. It does show that access to a fixed level of measured intelligence is losing scarcity unusually fast.
The question for every AI product is now concrete: if model capability becomes 13 times cheaper within a year, what value remains unique? Products with trusted context and dependable workflows have an answer. Services selling little beyond model access still need one.



