top of page

Meituan LongCat Owl Alpha Becomes OpenRouter Most Popular Model

Meituan LongCat Owl Alpha climbed to the top of OpenRouter rankings within days of release. The 1.6 trillion parameter mixture-of-experts model posted daily call volumes in the global top three and first place on the Hermes Agent leaderboard.

The release surprised observers because the entire training run used 50,000 domestic ASICs from manufacturers such as Biren Technology rather than established GPU clusters. Cumulative token consumption crossed 10 trillion within weeks, matching usage patterns previously seen only from Gemini and Opus 4.6 class systems.

Training Run Completed on Domestic Hardware

Meituan stated the model finished training on 35 trillion tokens using only chips designed and fabricated inside China. "This demonstrates our ability to scale frontier training entirely on domestic silicon," a company spokesperson said. No foreign GPU vendor participated in the final compute stage. This marks one of the largest publicly reported MoE training jobs performed entirely on non-NVIDIA silicon.

Performance benchmarks released alongside the launch placed Owl Alpha at parity with Gemini and Opus 4.6 on agentic coding and multi-step reasoning suites, per OpenRouter's official announcement and the Hermes Agent leaderboard. Independent testers at The Verge noted strong results on the Hermes Agent and Claude Code leaderboards, where it captured first and second place respectively.

The decision to train on domestic ASICs reduced reliance on export-controlled hardware. It also forced Meituan engineers to optimize kernels and memory layouts from scratch for the new accelerator design.

Daily Volume Surpassed Most Western Releases

OpenRouter data showed Owl Alpha holding a top-three spot in daily requests shortly after launch. It ranked third in the overall OpenClaw category while maintaining first position among agent-focused workloads. These figures appeared before any paid promotion on the platform.

Usage patterns indicated heavy adoption inside automated coding agents and research pipelines. Teams reported lower latency on multi-turn tool use compared with similarly sized dense models. The 10 trillion token consumption figure reflects both free-tier experiments and sustained production traffic.

Meituan confirmed the model would be retired in the coming weeks. The company plans to release follow-up versions without specifying architecture changes or training scale.

Domestic ASIC Supply Chain Faces New Test

The training run validated that Chinese accelerator designs can support frontier-scale MoE workloads when paired with sufficient software investment. Engineers worked around lower memory bandwidth by increasing expert parallelism and adjusting batch sizes during pre-training.

Critics point out that inference hardware availability outside Meituan remains limited. Most developers still access Owl Alpha through OpenRouter rather than local clusters. This creates a single point of distribution that could affect reproducibility once the model is taken offline.

Some industry observers question whether subsequent versions will retain the same training hardware or shift back to mixed GPU and ASIC fleets. Meituan has not released details on the next training run.

Agent Leaderboard Results Drive Adoption

The Hermes Agent ranking reflects performance on autonomous task completion sequences that require tool calling and long context retention. Owl Alpha led that category after only two weeks of public access. The Claude Code ranking measures code editing accuracy across multi-file repositories, where the model placed second.

These results matter because agent workloads now account for a growing share of paid API traffic. Developers prioritize models that maintain coherence across dozens of steps rather than single-turn chat performance alone. Owl Alpha demonstrated consistent behavior in both public benchmark runs and real OpenRouter traffic logs.

Next Milestones to Track

Three signals will clarify whether Meituan can repeat the result. First is the public release of the next LongCat model and whether it again trains on domestic ASICs only. Second is updated daily volume numbers on OpenRouter after Owl Alpha retirement to measure sustained demand. Third is any third-party inference cluster announcements that would let teams run the model outside the OpenRouter gateway.

Continued top-three placement after the current model is removed would indicate the training approach scaled successfully. A drop would suggest the initial surge depended on novelty and temporary capacity access.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

For better AI experience,

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

​Add Search Bar in Your Brain

Just Ask remio

Remember Everything

Organize Nothing

bottom of page