top of page

Xiaomi MiMo Tops OpenRouter as Chinese Models Sweep the Weekly Top Five

Aug 3
11 min read

Xiaomi’s MiMo-V2.5 reportedly processed 10.5 trillion tokens in one week, leading an OpenRouter ranking whose first five places went to Chinese models.

That sweep matters because OpenRouter measures routed usage, not benchmark victories or chatbot downloads. Developers use the platform to send requests across models and infrastructure providers through one interface.

The reported result therefore captures a specific market where switching is easy and consumption is measurable. It also puts pressure on premium models from OpenAI, Anthropic, and Google.

According to a weekly leaderboard report, MiMo-V2.5 grew 12 percent from the previous week. Two DeepSeek models occupied second and fifth place.

Tencent’s newly released Hunyuan model took third, while Zhipu AI’s GLM family reportedly placed fourth. Tencent’s model recorded the fastest reported growth after its recent release.

The result does not prove that Chinese models lead the entire global AI market. OpenRouter represents one large aggregation platform, not every direct API, enterprise contract, or consumer assistant.

It does show something narrower and still important. Chinese open-weight models have become default candidates for developers choosing models by workload, performance, and operating efficiency.

OpenRouter’s Top Five Changed Faster Than the Market Narrative

The week’s biggest change was not MiMo’s first-place finish alone. It was the concentration of every leading position among Chinese-developed models.

The reported ranking covered the week from July 20 through July 26, 2026. MiMo-V2.5 led with about 10.5 trillion routed tokens and 12 percent weekly growth.

DeepSeek V4 Flash followed with approximately 6.37 trillion tokens. Tencent’s Hunyuan Hy3 reportedly ranked third with about 3.94 trillion.

Zhipu AI’s GLM 5.2 held fourth place with roughly 3.29 trillion tokens. DeepSeek V4 Pro completed the top five, giving DeepSeek two distinct entries.

Those placements form more than a list of popular brands. They represent several product strategies competing within the same open-model market.

MiMo-V2.5 targets multimodal and agent-based workloads, which combine text, images, tools, and extended task state. Its model page describes a one-million-token context window.

OpenRouter lists the model as a native omnimodal system, meaning it handles several input formats within one architecture. The platform records an April 22, 2026 release date.

The MiMo model listing also emphasizes agent performance and long-context processing. Those features align with workloads that generate large token totals.

An automated coding agent can read repositories, revise files, call tools, and review its own output. A single task can consume far more tokens than a chatbot answer.

DeepSeek’s two positions reveal a different strategy. The Flash model serves frequent, efficiency-sensitive requests, while the Pro model targets more difficult coding and agent tasks.

That pairing lets developers route routine work and demanding reasoning to different models from the same family. It resembles a product line rather than one general-purpose endpoint.

Tencent’s rise adds another layer. The reported weekly increase exceeded 999 percent after Hunyuan Hy3 became openly available on July 6.

A percentage that large usually reflects movement from a small starting base. However, reaching the top three still requires substantial absolute traffic after that initial base effect.

The live model rankings change as new requests enter the measurement window. Any published snapshot should therefore include its dates and platform scope.

This distinction prevents a misleading conclusion. The figures describe tokens routed through OpenRouter, not all tokens processed by each company worldwide.

Direct traffic to vendor APIs remains outside this particular ranking. Private deployments and models running on customer infrastructure also remain largely invisible.

Even with those limits, the sweep carries weight. Developers using a neutral routing layer had five Chinese models at the top of their collective consumption list.

That is a sharper signal than a temporary social trend. Every routed token represents an application, experiment, agent, or user session choosing a model for actual work.

Why Developers Are Routing More Work to Chinese Models

Rapid adoption reflects an operational advantage: developers can test these models quickly, then keep the ones that deliver acceptable results at scale.

OpenRouter reduces the effort required to change models. Applications can maintain one integration while sending requests to several model families or hosting providers.

That structure weakens a traditional advantage held by large AI vendors. A developer does not need to rebuild an entire product whenever another model becomes competitive.

The platform’s historical data already showed this behavior before the latest sweep. Chinese open models rose from a very small share during late 2024.

OpenRouter’s usage study found that Chinese open-source models reached nearly 30 percent of platform token volume during some weeks in 2025.

Their average weekly share across the study period was about 13 percent. DeepSeek remained the largest open-model contributor, while Qwen and newer families expanded the field.

The report also found that capable open models can gain meaningful traffic within weeks. That pattern helps explain Tencent’s reported rise after the Hy3 release.

Frequent releases give developers more reasons to retest their routing decisions. Each launch can improve one capability, reduce latency, or support a longer context.

Open weights further widen distribution. They allow independent providers to host the same model, although licensing terms and technical requirements still vary.

This produces competition at two levels. Model developers compete over architecture and training, while infrastructure providers compete over availability and serving quality.

Users can benefit from that structure without operating their own clusters. They can select hosted versions while retaining more portability than a single-vendor application provides.

The strongest demand appears connected to coding and agent workflows. OpenRouter’s study found that programming requests grew from roughly 11 percent of tokens to more than half.

That measurement covered the study’s recent weeks, not the July 2026 leaderboard. Still, it explains why code-oriented models can accumulate enormous token totals.

Programming agents repeatedly inspect files, compare outputs, and revise plans. Long prompts also include repository context, error logs, documentation, and previous attempts.

One human request can trigger dozens of model calls. Token volume therefore increases faster than the number of people using an application.

Chinese models have been designed around this changing workload. DeepSeek emphasizes coding and reasoning, while MiMo combines long context with multimodal processing.

Tencent has also framed Hunyuan as a family spanning text, image, video, and three-dimensional content. That breadth supports internal products and external developer use.

The result is a dense release cycle across several well-funded companies. Xiaomi, Tencent, DeepSeek, and Zhipu AI do not rely on one shared organization or business model.

Their simultaneous presence reduces the chance that the ranking reflects one company’s isolated promotional campaign. Developers are spreading traffic across several Chinese model families.

However, low switching friction cuts both ways. The same users who adopted MiMo or Hy3 quickly can leave when another model performs better.

OpenRouter’s earlier analysis described the open-model market as increasingly fragmented. No individual open model consistently controlled more than roughly one-quarter of open-model traffic.

That makes a weekly lead valuable but fragile. Continued use matters more than the launch spike, especially after free access or introductory routing incentives end.

The Real Contest Is Open Routing Against Closed Model Lock-In

The central conflict is not China against the United States. It is portable model routing against dependence on one proprietary model stack.

National origin makes the top-five sweep politically striking. Yet developers usually make routing decisions through a more practical comparison.

They ask whether a model can complete the task reliably, fit the required context, call tools correctly, and operate within the application’s resource limits.

OpenRouter makes that comparison immediate. A team can assign one model to planning, another to implementation, and a third to final review.

That multi-model design directly challenges the idea that every AI application should inherit one vendor’s complete stack. It separates the application from the model beneath it.

OpenAI, Anthropic, and Google still hold major advantages. Their proprietary systems often lead difficult evaluations and offer mature enterprise controls, integrations, and support.

Anthropic has also held a strong position in programming-related spending, according to OpenRouter’s historical study. Claude remains a reference point for demanding coding work.

Google combines its Gemini models with cloud infrastructure, productivity software, and consumer distribution. OpenAI connects its models to a large developer base and ChatGPT.

Chinese open-weight models apply pressure somewhere different. They make acceptable performance available through more hosts and encourage application-level experimentation.

That pressure is strongest for workloads with high volume and repeatable evaluation. Coding pipelines can run tests, compare patches, and reject incorrect outputs automatically.

A business may still choose a premium proprietary model for a critical final decision. It can route document preparation or repetitive analysis through another model.

DeepSeek’s Flash and Pro pairing fits this pattern. Flash can handle frequent tasks, while Pro can address requests that need deeper reasoning.

MiMo’s long context creates another route. Developers can give one model more source material instead of building an elaborate retrieval chain for every request.

A long context window does not guarantee accurate recall. It simply raises the amount of material that an application can provide within one request.

The distinction matters for teams building knowledge-heavy agents. More context can reduce integration work, but irrelevant material can still confuse the model.

Good systems therefore need evaluation, traceability, and organized source material. A searchable engineering knowledge base can support those checks without determining which model must process the content.

The broader shift favors model-agnostic software. Such applications treat models as replaceable components while preserving data pipelines, permissions, and user experience.

This design limits vendor lock-in, but it introduces new engineering responsibilities. Teams must compare output quality and monitor changes across several providers.

They must also manage differences in tool schemas, safety behavior, context handling, and structured output. A shared interface does not make every model identical.

The OpenRouter sweep shows that many developers accept those costs. They are routing enough work to Chinese models for all five to lead the weekly ranking.

For proprietary vendors, the forced response is clear. They must justify restricted access through better reliability, stronger integrations, or measurable advantages on valuable tasks.

Brand recognition alone becomes less protective when developers can test another model in minutes. Closed vendors need results that survive direct workload comparisons.

Token Volume Is Not the Same as Revenue, Quality, or Trust

The leaderboard measures consumption on one platform, so it cannot independently establish technical leadership, commercial success, or enterprise acceptance.

Token totals reward workloads that produce long prompts and responses. Agent loops, roleplay sessions, repository analysis, and synthetic data generation can dominate raw volume.

One trillion tokens from automated tasks do not equal one trillion tokens from regulated business decisions. The value and risk attached to those requests differ.

Free access can also distort rankings. A newly released model can attract experiments because trying it requires little commitment.

The reported increase above 999 percent for Hunyuan Hy3 illustrates this problem. The growth is notable, but the starting point was necessarily much smaller.

A durable shift requires the model to retain traffic after launch attention fades. Weekly totals should remain elevated across several rolling windows.

Another limitation involves concentration. One large customer or application can materially affect a platform-wide chart.

OpenRouter acknowledged this effect in its historical analysis. One sizable account produced a noticeable spike in tool-related volume during May 2025.

That example does not invalidate the current ranking. It shows why observers should avoid converting a token leaderboard directly into global market share.

The most accurate claim is narrower. Chinese-developed models occupied the five highest positions in one reported weekly OpenRouter snapshot.

Quality is equally difficult to infer. Developers may select a model because it is sufficient for a task, not because it produces the best output available.

A model that handles routine classification can generate more tokens than a specialist used for difficult scientific analysis. Volume and capability answer different questions.

Company benchmark claims require similar caution. DeepSeek reportedly positions its flagship model against leading closed systems for code and complex agent tasks.

Those statements need independent evaluation across real repositories, tool failures, and long-running workflows. A benchmark score cannot capture every production constraint.

MiMo’s listed context length also needs practical testing. Models can accept long inputs while failing to retrieve a crucial detail buried within them.

Latency, rate limits, and provider reliability influence adoption too. An application needs predictable performance during peak demand, not only a strong model architecture.

Security remains a separate decision. Some organizations restrict particular model providers or hosting regions because prompts can contain confidential material.

Open weights can reduce that concern when customers deploy models within controlled infrastructure. However, using an open model through a third-party API still creates a data relationship.

Developers must examine the actual provider, retention policy, logging behavior, and geographic processing path. A model’s country of origin does not answer those questions alone.

Licensing also deserves attention. “Open source” is often used loosely for models whose weights are available under specific restrictions.

OpenRouter itself abbreviates open-weight models as OSS for simplicity. Buyers should still review the license before distributing or modifying a model.

Regulated enterprises may continue favoring proprietary systems with established compliance documentation. That choice can persist even when an open model leads token volume.

Consumer reach also remains missing. ChatGPT, Gemini, and other first-party products generate significant activity outside OpenRouter.

Direct enterprise agreements create another blind spot. Large companies often access models through cloud contracts rather than public aggregation platforms.

The ranking therefore describes the contestable layer of the market. It shows what happens where developers can compare models with relatively low switching costs.

That layer matters because it can influence future purchasing. A model first tested through an aggregator can later move into direct hosting or a private deployment.

Still, the July sweep should be treated as an adoption signal, not a final verdict. Its strongest message concerns developer willingness to switch.

What Xiaomi, DeepSeek, and Tencent Must Prove Next

Three signals will determine whether the top-five sweep marks a durable change or another fast-moving leaderboard cycle.

The first signal is retention. MiMo-V2.5, DeepSeek V4 Flash, Hunyuan Hy3, GLM 5.2, and DeepSeek V4 Pro must hold substantial traffic beyond launch windows.

Rolling four-week and quarterly totals will provide a better view than one weekly snapshot. Stable traffic would suggest integration into persistent applications.

A sharp decline would point toward experimentation, temporary access conditions, or migration to newer models. Open routing makes both adoption and departure unusually fast.

MiMo deserves particular attention because it led by a wide reported margin. Its position will strengthen if developers keep using its multimodal and long-context features.

The second signal is workload quality. Independent evaluations should test these models inside complete coding and agent systems, not only isolated benchmark prompts.

Useful tests should measure successful tool calls, completed repositories, error recovery, and performance across long tasks. They should also disclose infrastructure and evaluation methods.

DeepSeek V4 Pro faces the clearest test here. Its reported positioning depends on handling complex agent work near the level of leading closed models.

If those results reproduce across independent teams, proprietary vendors will face direct pressure on high-value workloads. If they do not, volume will remain an efficiency story.

Tencent’s Hy3 needs the same scrutiny after its rapid debut. A launch-driven rise becomes more meaningful when developers publish repeatable production results.

The third signal is the response from OpenAI, Anthropic, and Google. Their reaction will show whether they consider open-model routing a strategic threat.

A response can take several forms without focusing on public price figures. Vendors can improve smaller models, expand context, strengthen tool use, or loosen deployment options.

They can also deepen enterprise integration, where compliance and support remain stronger defenses. Better reliability may preserve premium demand even when raw token share shifts.

OpenRouter’s own data supports this divided market. Its study found open models concentrated in lower-cost, high-volume usage while closed systems retained valuable workloads.

The latest sweep suggests that Chinese developers are moving upward within the high-volume side. The unresolved question is how much premium work follows them.

Developers should therefore avoid choosing a winner from one chart. They should test models against their own tasks and preserve the ability to reroute.

A practical evaluation starts with representative prompts, expected outputs, and failure conditions. It should include latency, tool completion, context retrieval, and human review.

Teams should also record which model produced each output. That trace becomes essential when a provider updates a model or routing configuration.

Knowledge workers face a related issue. As applications mix several models, they need a stable place for source documents and verified decisions.

A personal knowledge system can preserve that continuity while the models underneath an application keep changing.

The reported OpenRouter sweep is important because it shows developers acting on choice. Chinese models did not merely enter the comparison set.

They reportedly occupied every leading position within a major multi-model marketplace for one week. Xiaomi, DeepSeek, Tencent, and Zhipu AI each contributed to that result.

Yet the next phase will be harder than climbing a live chart. Those companies must convert experimentation into retained workloads, trusted deployment, and repeatable performance.

OpenAI, Anthropic, and Google still possess distribution, enterprise relationships, and strong proprietary models. They are pressured, but they have not disappeared from global AI use.

The enduring change is the weakening of automatic model loyalty. Developers increasingly treat an AI model as a component that must keep earning its route.

Watch the next several weekly windows, independent agent evaluations, and closed-vendor responses. Together, those signals will reveal whether this sweep becomes a lasting market shift.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page