Why Firms Are Struggling to Set Prices for AI
- Aisha Washington

- 1 day ago
- 14 min read
Google News surfaced a BBC report on August 4 that captures a growing conflict in artificial intelligence: firms still cannot decide what an AI product is worth.
That uncertainty now affects model developers, software vendors, enterprise buyers, and the finance teams approving deployments. AI can be sold by user, token, action, conversation, capacity, or business result. Each choice transfers cost and risk differently.
The problem became harder as AI moved beyond chat. OpenAI and Anthropic popularized token-based access, while Microsoft and Salesforce brought agents into established business software. Agents can run several model calls, search company data, use tools, and revise their own work before completing one request.
Traditional software pricing assumes that adding another user creates little additional operating cost. AI breaks that assumption because every generated response consumes computing resources. Agentic systems make the mismatch larger by doing an unpredictable amount of work for each task.
The result is a contest between predictable subscriptions and variable consumption. Vendors need revenue that follows their computing costs. Customers need bills they can forecast and value they can defend. Neither side wants to absorb all the uncertainty.
The Google News Story Points to a Market Still Testing Its Business Model
The important change is not one new price. It is the industry’s retreat from a single, dependable way to sell software.
Software companies spent decades making subscriptions easy to understand. A business counted employees, selected features, and negotiated a contract. Usage could vary without changing the bill dramatically.
Generative AI introduced a direct link between customer activity and vendor expense. A token, the small unit used to process and generate text, represents work performed by a model. Longer prompts, larger documents, deeper reasoning, and repeated tool calls all increase consumption.
That system appears precise, but precision does not guarantee clarity. Most employees do not know how many tokens a task requires. A procurement team cannot easily translate tokens into completed reports, resolved cases, or faster product decisions.
The challenge grows when an agent handles the task. An ordinary chatbot might answer once. An agent can plan, retrieve records, call external systems, evaluate a result, and try again. One visible request may therefore create many hidden model operations.
OpenAI’s enterprise options illustrate this distinction. Its capacity model lets API customers reserve token throughput for a particular model snapshot. That helps manage availability, but buyers must still estimate how much capacity their applications need.
Microsoft uses another layer of abstraction. Selected Copilot and agent capabilities use consumption billing, while prepaid capacity remains available for some deployments. Its documentation says reasoning models can activate more than one billing meter during a single interaction.
These approaches are not inherently unfair. They reflect a service with variable production costs. However, they push forecasting work onto customers before those customers fully understand their future usage.
The BBC headline circulating through Google News matters because it describes a structural problem, not a temporary promotion cycle. AI vendors are testing how to connect computing cost, customer value, and understandable contracts. Those three elements rarely move together.
A cheaper model can reduce the expense of one call. It does not automatically make a multi-step workflow predictable. An efficient agent can complete work with fewer calls, while a poorly designed one can consume more resources without producing a better answer.
Pricing therefore becomes a product design issue. The meter influences how developers build workflows, how often users rely on them, and whether managers permit autonomous operation. It also shapes which failures become financially painful.
A fixed subscription encourages experimentation but exposes the vendor to heavy users. Token billing protects the vendor but can discourage adoption. Outcome billing sounds aligned with value, yet requires both sides to agree on what success means.
No model removes risk. It only decides who carries it.
Seat-Based Software Collides With Variable AI Costs
AI vendors want subscription simplicity, but the economics of inference keep pulling them toward consumption billing.
Seat-based pricing charges for each authorized user. It works well when software serves as a tool that employees operate directly. A customer can estimate spending by counting users, while the vendor benefits from recurring revenue.
AI agents complicate that logic because they can perform work without continuous human input. One employee might launch an agent that completes a single lookup. Another might start a workflow that reviews hundreds of documents and revises its output several times.
Charging both users equally can separate revenue from operating cost. It also separates price from delivered work. The customer running the larger workload receives more service even though both seats appear identical.
The opposite model charges for consumption. Vendors can meter tokens, messages, model calls, actions, or processing capacity. Revenue then follows technical activity more closely.
Microsoft’s agent billing guidance shows how detailed this can become. Different agent features consume different amounts, and reasoning-capable models can engage separate meters. Such granularity supports cost recovery, but it makes budgeting dependent on workflow behavior.
Customers can monitor usage, set limits, and optimize prompts. Those controls help after deployment. They do not solve the earlier question of how much an untested workflow will consume at scale.
This is especially difficult for knowledge work. A repetitive classification task has relatively stable inputs and outputs. Research, coding, contract analysis, and customer support contain exceptions that send an agent down longer paths.
A support agent might solve a familiar request with one retrieval. A rare account problem might require several database queries, policy checks, and handoffs. Billing based on activity captures that difference, but the customer pays more precisely when the system encounters difficult cases.
That can create a bad incentive. A vendor may earn more when an agent works inefficiently, while a customer bears the additional cost. The vendor can counter this concern with transparent logs, efficiency targets, and spending controls. Still, trust becomes part of the commercial agreement.
Fixed subscriptions reverse the incentive. The vendor benefits when models and workflows become more efficient because the revenue remains stable while computing expense falls. Yet unlimited access creates exposure when users automate large jobs or continuously run agents.
Hybrid pricing attempts to divide the risk. A company might include ordinary use within a license, then charge separately for advanced actions or higher consumption. This preserves a familiar entry point while limiting open-ended vendor costs.
Hybrid structures bring their own complexity. Buyers must understand which activities are included and which activate a new meter. If the boundary appears only after deployment, a simple subscription becomes an unpredictable variable contract.
The conflict pressures established software companies most. They trained customers to expect stable per-user contracts. Now they must introduce variable billing without making their products feel like unfamiliar cloud infrastructure.
Model providers face a different problem. Their unit of production is easier to measure, but their services increasingly compete with cheaper models. Customers can route straightforward tasks to smaller systems and reserve premium models for demanding work.
That routing weakens a vendor’s ability to charge a broad premium. It also makes the application layer more important because orchestration determines which model performs each task.
The Google News report is therefore not just about selecting a number. Firms must decide whether they are selling access, computing, labor, or a result. The answer changes the meter, the customer relationship, and the product itself.
AI Outcome Pricing Sounds Fair Until Someone Defines Success
Outcome-based pricing aligns payment with value in theory, but attribution turns the apparent solution into another negotiation.
An outcome model charges when the system produces an agreed result. A customer-service vendor might connect billing to resolved conversations. A sales product might use completed work or another verified event.
This approach has immediate appeal. Customers do not want to buy tokens for their own sake. They want fewer unresolved requests, faster analysis, better service, or completed business processes.
Vendors also gain a clearer value story. Instead of defending an invisible unit of computing, they can connect payment to something a manager already measures.
Salesforce now offers several Agentforce structures rather than one universal answer. Its Agentforce options include conversation-based billing and a flexible credit system. Salesforce introduced the latter after its initial approach proved insufficient for every workflow.
That evolution demonstrates why outcome pricing becomes difficult. A conversation can be short or long. An action can be trivial or consequential. A resolution can result from the agent, the customer’s own effort, a human employee, or several systems working together.
The parties must first define the billable event. They then need rules for duplicate requests, reopened cases, partial completions, customer abandonment, human intervention, and incorrect results.
Quality adds another layer. An agent might close a support case while leaving the customer dissatisfied. It might classify a lead as qualified without creating revenue. It might draft a contract that saves time but still requires extensive legal review.
If billing follows the first visible completion event, the vendor can optimize for closure rather than durable value. If payment depends on a later business result, outside factors can overwhelm the agent’s contribution.
Outcome pricing also changes accountability. A vendor accepting payment only for success appears to take more performance risk. In practice, contracts may narrow the definition of success, exclude uncertain cases, or require customers to maintain specific data and processes.
Those conditions can be reasonable. An agent cannot deliver reliable results from missing records, conflicting policies, or inaccessible systems. Still, every condition weakens the claim that the price simply follows value.
An independent buyer must ask who controls the outcome data. When the vendor defines the meter, operates the agent, and reports success, customers need auditability. They should be able to inspect the event that triggered billing and challenge incorrect classifications.
A second question concerns optimization. Does the agent stop when it reaches the billable threshold, or when it reaches the customer’s actual goal? Those two moments are not always identical.
A third question concerns failure. An agent can consume substantial computing resources without completing a task. Under outcome pricing, the vendor absorbs that direct cost. The likely response is to restrict difficult workflows, price uncertainty into the contract, or route borderline cases to humans.
This means outcome billing does not erase technical cost. It hides that cost behind a business event and moves the risk into eligibility rules.
For tightly defined, high-volume workflows, the model can work. Both sides can measure the event, examine exceptions, and estimate frequency. The arrangement becomes harder for open-ended knowledge work where quality is subjective.
A research memo, product strategy, or software design rarely has one binary success point. Its value emerges later through human decisions. Charging per outcome would require a disputed judgment about quality or impact.
That limitation keeps usage and seat pricing alive. They may be imperfect, but they meter observable things. Outcome models work best where the result is equally observable and attributable.
Cheaper Models Raise Pressure Without Solving the Forecasting Gap
Competition can lower the cost of intelligence, but lower unit costs do not make autonomous workloads predictable.
Open models and smaller specialized systems give companies more bargaining power. A development team can use a premium model for hard reasoning, then direct routine classification or extraction to a less expensive option.
This multi-model approach reduces dependence on a single provider. It also turns model selection into an operational decision rather than a permanent commitment.
The shift pressures OpenAI, Anthropic, Google, Microsoft, and other vendors to justify premium services. Raw benchmark leadership matters less when a smaller model completes the customer’s actual task reliably.
IDC argues that the AI contest is moving toward measurable results. Its outcome analysis says operational readiness remains a major constraint when companies struggle to move pilots into core workflows.
That distinction is crucial. A lower token rate helps only if the workload, prompt, retrieval system, and tool calls remain effective. A cheap answer that requires repeated correction can cost more than a strong first attempt.
Agent design determines much of the final bill. Developers choose how much context to provide, when to retrieve documents, which tools to call, and how many retries to allow. They also decide whether a task needs an advanced reasoning model.
Prompt caching offers one example of architectural savings. It lets an application reuse previously processed prompt content instead of recomputing the same material. OpenAI’s caching documentation describes how repeated prompt prefixes can receive different treatment from uncached input.
Caching helps when an application repeatedly sends stable instructions or reference material. It helps less when every task uses different records or requires fresh context.
Retrieval can reduce the amount of information sent to a model, but poor retrieval causes other costs. If the system selects irrelevant documents, the model may produce a weak answer or need another attempt.
Human review must also enter the calculation. An AI system can appear inexpensive at the API layer while shifting verification work to employees. A useful cost measure includes setup, monitoring, correction, security, and governance.
This is why comparisons based only on tokens can mislead buyers. Tokens are a production measure, not a complete measure of useful work.
The same problem affects subscription comparisons. A nominally unlimited plan can include rate limits, model restrictions, or policies that change how much heavy users receive. Businesses need service guarantees and workload tests, not only a plan label.
Google News coverage has increasingly reflected this tension between falling model costs and rising aggregate use. As agents perform longer tasks, efficiency improvements can encourage more consumption. Lower cost per step does not guarantee lower total spending.
This pattern resembles cloud computing. Cheaper storage and processing expanded what companies built, while overall cloud bills still required active management. AI introduces additional uncertainty because model behavior and workflow length are probabilistic.
A deterministic program follows a defined sequence. An agent can choose different paths for similar requests. That flexibility creates value, but it also complicates capacity planning.
Companies can respond with task budgets. An agent receives a maximum number of steps, tool calls, or tokens. It must stop, ask for approval, or hand the work to a person when it reaches the limit.
They can also route tasks by complexity. A lightweight model handles ordinary work, while a more capable system receives only difficult cases. Evaluation data should determine those routing rules.
For knowledge-intensive work, maintaining reliable context is equally important. A well-organized knowledge workflow can reduce unnecessary searching and repeated document processing. However, information quality still requires testing within the actual application.
The competitive outcome will not belong automatically to the cheapest model. It will favor systems that translate variable intelligence costs into controlled, dependable work.
What AI Pricing Metrics Still Fail to Show
Every current pricing model leaves out part of the value chain, so buyers should distrust claims that one meter fully aligns incentives.
Token pricing measures model activity. It does not measure whether the answer is accurate, useful, or necessary. An application can consume fewer tokens and still fail its task.
Seat pricing measures authorized access. It does not reveal how much work the system performs or whether employees adopt it. A company can license many users while receiving little operational value.
Action pricing measures steps taken. It can reward a system for doing more work, even when a shorter path would be better. The definition of an action may also vary across products.
Conversation pricing creates a recognizable customer-service unit. Yet conversations differ in length, complexity, and outcome. A reopened case can expose ambiguity about whether the original interaction succeeded.
Outcome pricing measures a declared result. It struggles with attribution, quality, delayed effects, and external factors. It also invites disagreement over who controls the measurement.
No meter captures everything. A buyer therefore needs a collection of technical and business measures.
The technical side should include consumption by workflow, model, environment, and task type. Teams need failure rates, retry counts, latency, tool calls, and human escalation.
The business side should include completion quality, time saved, adoption, customer response, and the cost of oversight. These measures should connect to a defined baseline from before the AI deployment.
Without a baseline, vendors and customers can both claim success. The vendor points to activity. The customer points to unchanged business results. Neither can establish what the system actually improved.
A controlled pilot should answer more than whether the agent can perform a task. It should show the distribution of consumption across easy and difficult cases.
Averages conceal dangerous variation. An agent might be economical on most requests but extremely expensive on a small group of exceptions. Those exceptions can dominate total spending after deployment.
Buyers also need to test adversarial and malformed inputs. An agent that enters a loop, repeatedly calls a tool, or processes unexpectedly large documents can consume resources without delivering value.
Spending limits are essential, but blunt limits can interrupt business processes. Teams should combine global limits with workflow-specific controls and alerts.
Governance matters because employees often lack visibility into the commercial effect of their actions. A user sees one button. The system behind it might invoke several models and enterprise services.
Clear interfaces should disclose when a task uses premium reasoning, large context, or an autonomous workflow. The goal is not to burden every user with token accounting. It is to help them understand consequential choices.
Vendors should expose billing data at the level customers can act upon. A monthly total is not enough. Teams need to identify which agent, workflow, or feature caused a change.
Customers should also resist false precision. A detailed credit system can appear transparent while remaining difficult to connect to real computing activity. Credits help package technical complexity, but only if their conversion rules remain stable and documented.
The skeptical view is that AI pricing will remain unsettled because the underlying products remain unsettled. Model capabilities change, inference methods improve, and agents take on new tasks. A durable meter cannot easily emerge while the unit of value keeps moving.
That does not make enterprise adoption impossible. It means contracts should preserve flexibility. Buyers need the ability to monitor consumption, change models, revise limits, and reconsider pricing as workflows mature.
Vendors, meanwhile, must avoid using complexity as cover. If customers repeatedly receive surprise bills or cannot explain charges internally, adoption will slow regardless of model quality.
Trust will depend on whether a company can predict spending before scale and reconcile it afterward.
Three Signals Will Show Where AI Pricing Goes Next
The winning pricing model will be the one that makes agent costs predictable without disconnecting payment from useful work.
The first signal is whether major software vendors simplify their meters after real enterprise deployments. Salesforce already offers per-user, conversation, and flexible consumption options. Microsoft combines licensing, prepaid capacity, and pay-as-you-go structures across its agent products.
More choices can support different workloads. They can also signal that vendors have not found one stable unit of value.
Watch whether those companies consolidate their options or add further distinctions. Consolidation would suggest that buyers and vendors have identified repeatable patterns. More layers would show that agent behavior remains too varied for a common contract.
The second signal is the quality of usage controls. Billing dashboards should move from monthly reporting toward workflow-level forecasting, automatic anomaly detection, and enforceable task budgets.
This matters more than another nominal price reduction. A finance team can manage a relatively expensive service when spending is explainable. It will hesitate over a cheaper service with unpredictable exposure.
Better controls would strengthen consumption pricing. Customers may accept variable bills when they can trace, forecast, and cap activity. Weak controls would push buyers back toward fixed subscriptions or tightly scoped pilots.
The third signal is whether outcome pricing survives contact with messy business processes. Customer support offers one of the clearest tests because conversations, resolutions, reopenings, and escalations can be recorded.
If vendors and customers agree on durable definitions, audit disputes efficiently, and preserve service quality, outcome billing can expand into other structured workflows.
If contracts accumulate exclusions and buyers dispute what counts as success, outcome pricing will remain a selective option rather than the default.
Independent analysis of enterprise deployments will be more useful than vendor announcements. Buyers should look for evidence covering total operational cost, not just model consumption. That includes integration, monitoring, human review, and failed work.
The next generation of agent products will likely support several commercial models. Routine employee assistance can fit within a user license. High-volume automation can use metered consumption. Narrow workflows with verifiable results can support outcome billing.
That mixed future is less elegant than one universal answer. It is also more realistic because AI products perform different kinds of work.
For developers, pricing architecture now belongs in system design. Model routing, caching, context management, retry policies, and approval gates all affect the commercial product.
For enterprise buyers, procurement can no longer finish before implementation begins. Contract terms must reflect observed workload behavior, and technical teams need access to billing data.
For knowledge workers, the question is not whether every prompt has a visible charge. It is whether organizations restrict useful tools after unexpected consumption appears.
Google News has highlighted a genuine fault line in the AI economy. Vendors are selling software whose operating costs resemble infrastructure and whose promised value resembles labor. Neither traditional pricing model fits cleanly.
The decisive question is practical: can your organization connect each AI workload to a controlled cost and a measurable result? Until vendors make that answer easier, the safest approach is limited deployment, transparent metering, and expansion only after the economics survive real use.


