top of page

TRM Labs AI Agent Payments Research Finds x402’s Autonomy Story Is Ahead of Reality

Sep 14
12 min read

TRM Labs analyzed 198.9 million x402 settlements and found a sharp conflict between payment activity and actual AI autonomy. Its AI agent payments research estimates that agents generated only 0.6% to 7.5% of screened commerce by value.

That range challenges one of agentic commerce’s most attractive narratives. A blockchain can show that software sent a payment, but it cannot show whether an AI model chose the merchant, evaluated alternatives, or acted independently.

The distinction matters because Coinbase, Amazon, Google, OpenAI, Stripe, and major payment networks are building infrastructure around machine-initiated transactions. Their rails increasingly work. The evidence that autonomous agents are using those rails at meaningful commercial scale remains much thinner.

TRM’s findings do not mean agentic commerce has failed. They show that infrastructure adoption, automated payment volume, and intelligent economic activity are three different measurements. Treating them as interchangeable creates a distorted picture of the market.

TRM Labs AI Agent Payments Research Cuts the Headline Volume in Half

TRM’s most important finding is not that x402 lacks activity. It is that much of the activity cannot be treated as ordinary commerce.

The blockchain intelligence company examined known x402 facilitators across Base, Solana, and Polygon. These facilitators verify signed payment authorizations, broadcast settlements, and cover the blockchain network fee.

The dataset contained roughly USD 52.7 million across 198.9 million settlements since May 2025. At first glance, those figures describe an established machine-payment network with considerable transaction frequency.

TRM then filtered the data to remove activity that looked unlike genuine commerce. Its screens excluded addresses paying themselves, bulk flows dominated by one or two payers, and sellers with fewer than ten distinct buyers.

Those exclusions reduced the likely commerce pool to USD 25.62 million. In other words, about half of the observed value disappeared before TRM even tried to decide whether an AI agent was involved.

The full onchain analysis explains why this first filter is necessary. Scheduled processes, load tests, self-dealing, token activity, and ordinary software can all generate valid x402 settlements.

A technically successful payment therefore answers only a narrow question. It confirms that an address authorized a transfer under the protocol’s rules. It does not establish who controlled the address or why the payment occurred.

TRM faced a harder attribution problem after identifying likely commerce. An AI agent and a conventional script can follow exactly the same payment sequence. Their transactions can look identical on a public blockchain.

The company addressed that uncertainty with two models. Its permissive model counted payments that were facilitator-broadcast, variable in value, and below one dollar on average.

The strict model required those patterns to continue across multiple months. It also required either a public agent registration or payments to more than one seller.

Applied to the screened commerce pool, those tests produced the 0.6% to 7.5% estimate. The wide interval reflects uncertainty in the classification method, not a precise census of every agent.

That caveat is central to the story. TRM is not claiming that it identified every autonomous payment. It is estimating a plausible range from behavioral signals that remain incomplete.

The study also found that USDC dominated settlement. About USD 52.47 million of USD 52.68 million, or 99.6%, used that stablecoin across the unrestricted asset analysis.

This concentration makes x402 activity easier to describe as a stablecoin payment market. It does not make the underlying buyers easier to classify as intelligent agents.

The result changes the meaning of x402’s headline numbers. They document demand for programmable payments, but they do not independently document demand from autonomous AI buyers.

Why x402 Activity Does Not Prove Agentic Commerce

x402 can make a payment machine-readable without making the paying machine intelligent.

The protocol adapts HTTP status code 402, labeled “Payment Required,” into a working payment step. A client requests a digital resource, and the server responds with payment terms.

The client signs an authorization and repeats the request with that proof attached. A facilitator verifies the authorization, settles the payment onchain, and returns evidence that the resource can be delivered.

This design removes the conventional checkout page. It lets software purchase an API response, data query, model inference, or other digital service inside a normal request flow.

That mechanism is useful for AI agents because it keeps payment inside their execution loop. An agent does not need to open a browser, complete a card form, or ask a human to approve every tiny purchase.

However, none of those steps requires reasoning. A developer can write deterministic software that requests the same endpoint on a schedule, pays the same amount, and stores the response.

A load-testing system can do the same thing at much higher frequency. A service operator can also generate self-payments or promotional activity that increases transaction counts without representing external demand.

TRM’s classification uses price variation as one sign of agency. The reasoning is that an agent exploring several services should encounter different prices, while a script repeatedly calling one endpoint often pays a fixed amount.

This assumption is defensible, but it creates a blind spot. A legitimate single-purpose agent might repeatedly buy the same service at the same price. TRM’s model could classify that activity as conventional automation.

The opposite error is also possible. A script can rotate among endpoints, vary transaction amounts, or imitate an exploratory payment pattern without using an AI model.

Blockchain data alone cannot resolve that problem. The chain records addresses, assets, amounts, timestamps, and contract interactions. It does not record a reliable explanation of the decision process behind each authorization.

Public agent registries offer another signal, but registration remains voluntary. A wallet owner can declare that an address belongs to an agent, yet the declaration may not be independently verified.

TRM said most participants do not currently use these registries. That leaves researchers without a consistent identity layer connecting wallets to agents, operators, models, or delegated users.

The analytical problem resembles automated web traffic. A server can observe that a machine requested a page, but additional evidence is needed to distinguish a search crawler, a malicious bot, a testing tool, and an AI assistant.

Payments raise the stakes because classification affects fraud monitoring, sanctions screening, disputes, and responsibility. A mistaken label is more consequential when software can move funds at machine speed.

The x402 findings therefore expose a measurement gap. The market has reliable settlement data, but it lacks equally reliable attribution data.

That gap explains why transaction totals should not serve as a substitute for adoption metrics. Real agentic commerce requires evidence of delegated authority, goal-directed selection, and an economically meaningful purchase.

The Real Opponent Is the Agentic Commerce Promise

The primary conflict is not x402 versus card networks. It is the promise of autonomous commerce versus evidence of ordinary automation.

Technology companies have spent more than a year assembling the pieces needed for agents to shop and pay. Those investments are real, even if current transaction data does not prove broad adoption.

Google introduced its Agent Payments Protocol, or AP2, in September 2025. AP2 provides a shared framework for communicating user authority, purchase intent, and accountability across agents, merchants, and payment providers.

The AP2 framework launched with support from more than 60 technology, retail, and financial organizations. It also included an x402 extension for crypto-based agent payments.

OpenAI and Stripe took a consumer-facing route with the Agentic Commerce Protocol. Their initial Instant Checkout experience kept a person involved in confirming the purchase, even though ChatGPT handled product discovery and communication with the merchant.

That checkout protocol illustrates an important difference. AI-assisted commerce can be commercially relevant without giving a model complete spending autonomy.

The user still confirms the item, shipping information, and payment details. The merchant remains responsible for accepting the order, processing payment, fulfilling it, and supporting the customer.

Amazon has pushed further into autonomous software purchasing. Amazon Bedrock AgentCore payments entered preview in May 2026 and became generally available in August.

The service lets developers connect supported wallets, set spending limits, and allow agents to pay for APIs, Model Context Protocol servers, content, and other digital resources. Model Context Protocol, or MCP, is a standard for connecting AI systems with external tools and data.

According to the AgentCore release, the infrastructure can negotiate x402 payments, enforce limits, and record transaction activity. Those controls target enterprise concerns that raw wallet access cannot address alone.

These launches demonstrate strong supply-side commitment. Large companies are creating protocols, identity mechanisms, wallets, merchant connections, and governance controls before autonomous demand becomes easy to measure.

That sequence is common in infrastructure markets. Developers need reliable rails before they can build applications, while payment providers need credible applications before transaction volume becomes economically meaningful.

The risk comes from confusing readiness with adoption. A protocol can support autonomous purchasing even when most current users are scripts, tests, or narrowly programmed services.

TRM’s results put pressure on infrastructure providers to report better evidence. Transaction count and settled value no longer provide enough context when the same rail serves agents, ordinary automation, and anomalous activity.

Useful reporting would separate external commerce from self-generated flows. It would also distinguish AI-assisted purchases, deterministic machine payments, and goal-directed autonomous decisions.

That level of classification is difficult, but it is necessary. Otherwise, agentic commerce risks becoming a label applied to any software-generated payment.

The optimistic interpretation is that companies are building ahead of demand. The skeptical interpretation is that the market has developed sophisticated plumbing for buyers who rarely arrive.

Both readings fit the available facts. The next stage depends on whether developers deploy agents that discover and purchase resources across a meaningful range of independent sellers.

Agent Payments Explained by What the Models Cannot See

TRM’s range is valuable because it makes uncertainty visible, but its behavioral tests cannot provide definitive attribution.

The permissive and strict models create boundaries rather than a final answer. That approach is more credible than presenting every x402 settlement as agent activity.

Still, the lower and upper estimates depend on assumptions about how intelligent agents should behave. Agents that explore multiple sellers and encounter variable prices fit the model more easily.

Narrow agents fit less comfortably. Consider an automated research agent that buys the same specialized dataset every morning. It could evaluate current conditions, decide the purchase remains useful, and repeatedly pay one provider.

Its onchain pattern might resemble a scheduled script. The intelligent work could occur before the payment, while the settlement itself remains uniform.

A more elaborate script can create the reverse impression. It might select among several providers using fixed rules, produce variable payments, and remain active for months.

That system could satisfy several behavioral screens without using generative AI. Classification based on transactions would overstate autonomous agent participation.

TRM acknowledges this limitation. Its report says the price-variation test may undercount single-purpose agents that repeatedly pay one service.

The measurement window creates another constraint. x402 primarily supports machine-native payments for digital resources. It does not capture every card-based purchase initiated or assisted by an AI product.

Visa and Artemis describe a distinction between macro commerce and micro commerce. Macro commerce includes familiar purchases, such as travel or subscriptions, where an agent acts for a person through existing merchant systems.

Micro commerce covers small software-to-software payments for APIs, compute, or data. These payments can occur frequently and fall below the economics of conventional card transactions.

Their payment research reported adjusted x402 activity of roughly USD 15 million across 109.6 million transactions as of April 21, 2026. The different totals and dates show why methodology matters when comparing studies.

TRM used a later observation period and known facilitators across three networks. It also applied its own commerce and agency screens. The resulting figures should not be combined as if they measured an identical population.

Neither dataset captures every form of AI-supported shopping. A consumer might ask ChatGPT to compare products, then complete the purchase manually on a retailer’s website.

That transaction reflects AI influence, but not autonomous spending. Calling it agentic commerce would blur product discovery, recommendation, checkout assistance, and delegated purchasing.

The distinction matters to merchants. AI-generated referrals can affect which products customers consider, even when the final payment happens through a conventional checkout.

Payment infrastructure providers face a different question. They need to know when software has authority to initiate a transaction and who bears responsibility if the result violates the user’s instructions.

Researchers therefore need several metrics, not one. Useful categories include agent-assisted discovery, human-confirmed checkout, deterministic machine purchases, and autonomous goal-directed spending.

This classification would also improve comparisons across card and stablecoin systems. It would prevent infrastructure choices from being mistaken for different levels of intelligence.

Until those measurements exist, TRM’s range should be read as a bounded estimate of visible x402 behavior. It is not a verdict on every AI-mediated purchase across the internet.

Working Payment Rails Still Lack Agentic Accountability

The most urgent problem is no longer whether software can transmit money. It is whether every payment carries trustworthy evidence of authority and responsibility.

Traditional online commerce assumes that a person controls the buying session. Fraud systems, confirmation screens, receipts, disputes, and chargebacks all build on that assumption.

An autonomous agent changes the sequence. A person may grant a broad objective, while software chooses the seller, amount, timing, and resource without requesting approval for each transaction.

That arrangement creates several layers of responsibility. The user defines the objective, the agent platform executes it, the model interprets it, the wallet signs, and the merchant fulfills the request.

If the agent buys the wrong resource, each participant can point elsewhere. The user might blame the model, the platform might cite delegated authority, and the merchant might rely on a valid signature.

A malicious instruction can create a harder case. Prompt injection occurs when untrusted content manipulates an AI system into taking an unintended action.

An agent browsing external resources could encounter instructions designed to redirect a payment. A valid blockchain settlement would confirm execution, not legitimate intent.

The IMF has highlighted this tension between probabilistic AI decisions and deterministic payment systems. Its payments assessment identifies authorization traceability, cybersecurity, unclear liability, and correlated automated behavior as emerging risks.

Spending caps reduce exposure, but they do not establish why a transaction occurred. Logging improves investigation, but logs must connect the user’s mandate with the agent’s reasoning and the final payment.

Identity creates another challenge. A payment address can be pseudonymous, while a public registry entry may be voluntary or unverified.

TRM argues that agentic commerce needs accurate registration, counterparty reputation, and monitoring designed for high transaction counts. Conventional controls often focus on monetary value because human payments are less frequent.

Agent payments invert that pattern. One agent can generate many small transfers whose individual values remain below familiar review thresholds.

This volume-value mismatch can hide operational problems. Thousands of incorrect micropayments may create little immediate financial damage while revealing a compromised agent or defective policy.

The industry needs machine-readable mandates that describe what an agent can purchase, from whom, within which limits, and for how long. Those mandates must also be revocable.

Merchants need verifiable agent identities and evidence that the requested purchase fits the mandate. Operators need logs that preserve the connection between intent, decision, authorization, and settlement.

Users need understandable controls. A technical wallet signature offers little reassurance if the person cannot see what authority was granted or stop it quickly.

Teams adopting transacting agents should preserve their own evidence trails as well. Searchable records across prompts, approvals, and outputs can make complex AI workflows easier to review.

None of these controls proves that demand will grow. They make experimentation safer and measurement more credible.

The accountability layer will determine whether agent payments move beyond controlled developer environments. Without it, more settlement volume can create more ambiguity instead of more trust.

Three Signals Will Show Whether x402 Agent Payments Are Becoming Real

The next phase should be judged by verified agent behavior, independent merchant demand, and stronger attribution, not raw transaction totals.

The first signal is sustained spending across independent sellers. TRM’s stricter model already treats multi-seller activity as stronger evidence of agency because it suggests discovery and selection.

Future data should show whether registered agents repeatedly purchase from unrelated providers over several months. A broader merchant mix would strengthen the case that agents are making contextual choices.

Concentration would weaken it. If most value continues flowing through one payment contract, one router, or a small cluster of related addresses, infrastructure activity will remain difficult to separate from genuine market demand.

The second signal is adoption of verifiable mandates and agent identities. Google’s AP2, wallet providers, registries, and enterprise platforms approach this problem from different directions.

The decisive development will be a common record connecting the human or organization, the authorized agent, the permitted action, and the resulting settlement. Voluntary labels alone will not be enough.

A meaningful increase in authenticated transactions would strengthen TRM’s upper estimate. Continued low registration would preserve the current uncertainty, even if total volume rises.

The third signal is production evidence from large platforms. Amazon’s AgentCore payments reached general availability in August 2026, giving developers a managed route to transacting agents.

The relevant metric is not how many endpoints appear in a catalog. It is whether external agents buy useful resources under enforced budgets, with low error rates and repeat demand.

Published case studies should disclose the number of distinct outside buyers, seller concentration, repeat usage, failed transactions, and human intervention. Those figures would help separate functional demonstrations from durable commerce.

Consumer commerce needs similar clarity. A purchase that requires a person to press a confirmation button remains valuable, but it belongs in a different category from autonomous spending.

Developers and enterprise buyers should ask vendors which category their product supports. They should also ask how the provider distinguishes AI decisions from conventional automation.

For now, TRM Labs AI agent payments research supports a restrained conclusion. x402 has shown that software-native settlement can operate at high frequency across public blockchains.

It has not shown that autonomous AI agents account for most of that activity. The best available estimate places their share of screened x402 commerce between 0.6% and 7.5%.

That gap is not merely a branding dispute. It affects investment decisions, merchant priorities, security controls, and the credibility of an emerging market.

The rail works, but the buyer remains hard to identify. Over the next several months, watch for multi-merchant behavior, verifiable delegated authority, and independently reported production usage.

Those signals will reveal whether agentic commerce is moving from prepared infrastructure to real economic agency. Until then, every large transaction total deserves a second question: what evidence shows that an AI agent actually made the decision?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page