Zhipu AI’s 400% Revenue Surge Hides a Much More Expensive Race
Zhipu AI reported a roughly 400% increase in first-half revenue, but the Chinese model developer still spent more than twice its sales on research and development. The result is not a simple growth story. It is a test of whether surging model usage can become a durable business before infrastructure costs and aggressive competitors absorb the gains.
The Beijing company, now branded internationally as Z.ai, generated RMB953.9 million in the six months ended June 30, 2026. That was 399.7% above the same period last year. Revenue from its open platform and application programming interfaces, or APIs, reached RMB825 million and supplied 86.5% of the total.
That shift matters more than the headline percentage. Zhipu previously depended heavily on customized, locally deployed systems for enterprises and public-sector customers. It is now presenting itself as a cloud platform whose models can be called repeatedly by developers, software companies, and AI agents.
The emerging contest is between that scalable API model and Zhipu’s continuing need to finance frontier-model research. DeepSeek, Alibaba, Moonshot AI, and MiniMax are making the same customers difficult to retain. They release new models quickly, compete on price, and give developers more reasons to switch providers.
What Changed Behind Zhipu AI’s 400% Growth
The defining change was not simply higher revenue. Zhipu replaced project-heavy deployment work with metered model usage as its main business.
Zhipu published its interim results on August 31, covering the first six months of 2026. That establishes the underlying event behind the September 1 news cycle. The reporting period ended on June 30, so later usage and annualized revenue figures should not be confused with recognized first-half sales.
The company’s total revenue rose from approximately RMB191 million to RMB953.9 million. Its open-platform and API revenue increased 2,735.7%, from RMB29.1 million to RMB825 million. That category represented 86.5% of revenue, compared with 15.2% one year earlier.
API revenue comes from programmatic access to models. A developer sends text, code, images, or other inputs to a hosted model and pays according to usage or a related service agreement. This structure can scale without sending an engineering team to every customer site.
By contrast, Zhipu’s local deployment revenue declined 20.5% to approximately RMB129 million. Local deployment generally places a model and supporting infrastructure inside a customer-controlled environment. It can satisfy security and data-residency requirements, but installations tend to require more customization and longer sales cycles.
The revenue mix therefore moved sharply in one reporting period. During 2025, local deployments produced RMB534 million and represented 73.7% of full-year revenue. Six months later, cloud access had become the dominant line.
This is the strongest evidence that Zhipu is becoming a model platform rather than remaining an AI contractor. The distinction affects growth, margins, customer concentration, and how investors judge the company.
Zhipu also reported RMB251.6 million in gross profit, an increase of 163.7%. Gross margin, however, was 26.4%, below the rate implied by the prior-year numbers. Revenue expanded faster than gross profit because serving model calls still carries substantial computing costs.
The interim financial results also showed the scale of the spending behind the expansion. Research and development expenses rose 33.6% to RMB2.13 billion. The loss attributable to shareholders was about RMB2.07 billion, while the adjusted net loss reached RMB1.96 billion.
Those figures create the article’s central tension. Zhipu has found a much faster sales channel, but its cost structure has not yet turned that adoption into a self-financing model business.
API Demand Is Replacing Custom Projects
Cloud usage gives Zhipu a more repeatable source of revenue, but repeatability must be measured over several reporting periods.
A local AI project often produces uneven revenue. Contracts can be large, yet recognition depends on delivery schedules, testing, acceptance, and customer budgets. Winning another customer may require another deployment process.
An API business behaves differently. Once the platform exists, thousands of developers can call the same underlying service. Increased usage can generate more revenue without creating a separate implementation project for every customer.
Zhipu said its platform had more than 7.4 million users by the end of the first half. It also reported that token volume, a measure of the text and code processed by its models, had increased more than fortyfold since the beginning of 2026.
Tokens are small units used to process model inputs and outputs. Rising token volume indicates more activity, but it does not automatically indicate better economics. Free credits, discounts, lower prices, or compute-heavy workloads can increase usage without producing proportional gross profit.
The company attributes the growth to newer GLM models, coding demand, and infrastructure expansion. Its GLM family competes across general reasoning, software development, agents, and lower-cost inference.
Coding is particularly important because it can produce frequent, measurable calls. An AI coding assistant may query a model repeatedly while reviewing a repository, creating files, testing changes, and correcting errors. That pattern can generate more recurring consumption than an occasional chatbot session.
The company says its models have progressed from code completion toward agentic engineering. This term describes systems that plan and execute multiple software-development steps rather than suggesting a single line of code. The claim should be treated as Zhipu’s description of its product direction, not independent proof of reliability across production environments.
Zhipu also disclosed an annualized recurring revenue figure of $1.6 billion for August, up from $1 billion in early July. ARR annualizes a recent revenue run rate. It is a forward-looking operating indicator, not revenue already recorded under financial reporting rules.
The distinction is essential. First-half recognized revenue covered six completed months and totaled RMB953.9 million. The August ARR figure extrapolates a much shorter period. It can reveal acceleration, but it can also move quickly when usage, capacity, pricing, or promotional activity changes.
According to a detailed business mix analysis, management said August ARR rose 60% from the early-July level. That pace explains investor interest, but it makes the next reported period more important.
The market now needs evidence that API customers keep calling the models after initial launches and incentives. It also needs to see whether revenue growth can continue without gross margin deteriorating.
This is why the 400% headline is only the opening number. The business-model transition becomes credible when usage remains high, customer cohorts mature, and each additional unit of revenue requires proportionally less spending.
The Real Opponent Is Zhipu’s Cost Base
Zhipu is not only racing other laboratories. It is racing the cost of training, hosting, and improving the models that generate its revenue.
The company’s RMB2.13 billion research and development expense was more than twice its first-half revenue. That spending supported model training, inference capacity, engineering, and product development. It also shows why rapid top-line growth has not yet produced profitability.
Inference is the computing performed whenever a deployed model answers a request. A provider must supply processors, memory, networking, storage, and power for every call. More usage creates more sales, but it also increases variable costs.
Training creates another burden. Frontier models require large computing clusters before they can serve customers. A company must often commit that capital before knowing whether a model will attract sufficient paid demand.
Zhipu says infrastructure work reduced its cost per generated token by 80% from the beginning of 2026. It also says end-to-end service performance improved threefold and revenue associated with each unit of training and inference capacity increased fourteenfold.
These measurements are encouraging, but they are company-reported operating metrics. The public financial results do not provide enough detail to reproduce every calculation or separate improvements caused by engineering from those caused by higher utilization.
The company has been expanding inference capacity based on Chinese-designed processors. It says its platform now operates at a scale of approximately 100,000 domestic chips. That strategy addresses access constraints created by United States controls on advanced semiconductor exports.
Domestic hardware can reduce exposure to foreign supply restrictions. It can also require extensive optimization because software tools, networking systems, and model kernels have historically been developed around Nvidia’s ecosystem.
Zhipu must therefore solve two problems simultaneously. It needs enough capacity to prevent service bottlenecks, and it must make that capacity economical enough to improve gross margin.
The first problem has become visible across China’s AI market. Moonshot AI temporarily stopped accepting new Kimi subscriptions in July after demand exceeded available capacity. The episode, described in an AI capacity report, showed that a popular model can become operationally difficult to serve within days.
The second problem appears directly in Zhipu’s accounts. Gross profit grew much more slowly than revenue, while adjusted losses remained larger than sales. Zhipu has more demand, but it has not yet shown that scale will produce attractive unit economics.
Management’s efficiency figures imply that newer infrastructure is beginning to work harder. The financial test will arrive when those technical improvements flow into gross margin, operating cash requirements, and a smaller loss relative to revenue.
A cloud model provider can tolerate heavy investment while usage grows quickly. It cannot assume indefinite access to capital, especially when customers can move workloads among several compatible APIs.
DeepSeek, Alibaba, and MiniMax Keep Switching Costs Low
Zhipu’s API transition is happening inside a market where model loyalty remains weak and each new release can reset customer expectations.
Developers rarely select a model provider once and stop evaluating alternatives. Many applications route prompts between vendors based on cost, latency, model quality, availability, or data requirements.
OpenAI and Anthropic face this behavior globally. In China, the pressure is intensified by DeepSeek, Alibaba’s Qwen models, ByteDance’s Doubao family, Moonshot’s Kimi, and MiniMax.
DeepSeek established a reference point for lower-cost reasoning and open-weight distribution. Open-weight models allow developers to download model parameters under applicable license terms, then operate or modify the systems without relying exclusively on the original provider’s hosted endpoint.
Alibaba combines Qwen development with a large cloud business. That pairing gives it infrastructure, enterprise relationships, and multiple ways to bundle AI services. ByteDance can connect model development with enormous consumer distribution.
Moonshot has used large Kimi releases to attract developers interested in coding, agents, and long-context workloads. Its capacity interruption demonstrated strong demand, but it also highlighted the operational risk of sudden adoption.
MiniMax offers the clearest public financial comparison among China’s listed model specialists. The company said first-half 2026 revenue rose 283.1% to $116.6 million, exceeding its entire 2025 revenue.
MiniMax attributed the increase to paying users, enterprise customers, API calls, and its token subscription product. Its first-half results indicate that Zhipu is not alone in converting model interest into faster sales.
This comparison does not establish which company has the best technology or business. Their product mixes, reporting currencies, customer segments, and accounting policies differ. It shows that the market around Zhipu is expanding quickly enough to support several growth stories.
That makes retention more important than raw registration totals. A developer can create an account, test a model, consume promotional credits, and leave. A durable platform needs production workloads that continue generating paid calls after a competing model launches.
Reliability also matters. Enterprise buyers evaluate uptime, latency under load, security controls, auditability, geographic availability, and support. A benchmark win cannot compensate for an endpoint that slows down during a critical workflow.
Pricing remains another source of pressure. Chinese developers have repeatedly reduced model costs or offered discounted access to encourage adoption. Lower prices can expand the total market, but they can also prevent providers from converting technical efficiency into higher margins.
The AI token market has already produced changes in both directions. Some providers have competed aggressively on token prices, while cloud operators have raised charges for selected computing and storage services.
Zhipu sits between those forces. Its customers expect models to become cheaper, while the infrastructure beneath those models can become more expensive or harder to acquire.
The result is a market with limited pricing power. Zhipu’s recent growth gives it more usage data, customer feedback, and revenue to fund development. It does not give the company permanent protection from the next DeepSeek, Qwen, Kimi, or MiniMax release.
What the Numbers Still Do Not Prove
One exceptional half does not establish durable recurring revenue, strong margins, or a path to profitability.
The first uncertainty concerns the comparison base. Zhipu’s first-half 2025 revenue was only about RMB191 million. A small starting point makes a 400% increase mathematically easier than it would be for a mature cloud provider.
That does not make the growth unimportant. Adding more than RMB760 million of revenue in one year is substantial. It means the percentage should be evaluated alongside absolute revenue and the company’s spending.
The second uncertainty concerns the quality of API revenue. The reported total does not reveal how much came from long-term enterprise workloads, short-term demand around new models, prepaid arrangements, or concentrated customers.
Customer concentration can make growth less predictable. A small number of large buyers may generate enormous token volume, then renegotiate prices or move workloads when another model performs better.
The third uncertainty is the relationship between recognized revenue and ARR. Zhipu’s August ARR indicates a much faster run rate than its first-half accounts. That difference can reflect genuine acceleration after June, but readers should not treat annualized August activity as cash already earned.
ARR definitions also vary between companies. Some calculate the figure from contracts, while others annualize recent consumption. A usage-based business can produce an impressive run rate during a busy month without guaranteeing the same volume for twelve months.
The fourth uncertainty is profitability. Gross margin of 26.4% leaves limited room for research, sales, administration, and continued infrastructure development. Zhipu’s adjusted net loss was more than twice its revenue.
Losses did narrow on some measures, yet the company remains dependent on outside financing to sustain its current research program. The central question is not whether losses exist. Frontier-model developers commonly spend heavily. The question is whether losses shrink as a percentage of revenue when API usage scales.
The fifth uncertainty concerns capital intensity. Management presents the open platform as a scalable business, but serving model calls requires physical infrastructure. Growth can demand new chips and data-center capacity before customer revenue arrives.
The 80% reduction in token cost would help if measured consistently and sustained. However, the effect can be offset when newer models use more computation, customers demand longer context, or competitive pricing transfers efficiency gains to users.
The sixth uncertainty is model leadership. Quality rankings change rapidly, and no benchmark captures every production workload. Coding agents that perform well in controlled tests can still make expensive errors when they encounter private repositories, unfamiliar systems, or incomplete instructions.
Independent reporting on Zhipu’s results emphasizes the same duality. Revenue is growing quickly, while the company continues to carry substantial research costs and losses. The reported loss figures make it premature to describe the business-model shift as complete.
For enterprise buyers, this uncertainty supports a practical response. They should evaluate model quality with their own tasks, track total inference costs, and keep critical workflows portable where feasible.
For developers, the story is less about choosing a permanent winner. It is about understanding that fast competition can improve models and lower costs while increasing the risk of sudden pricing, capacity, or product changes.
Three Signals to Watch After the 400% Surge
The next phase will be decided by revenue quality, margin movement, and Zhipu’s ability to retain workloads after competing releases.
The first signal is Zhipu’s next recognized API revenue figure. Investors should compare it with the RMB825 million reported for the first half, rather than relying only on August ARR.
A strong second half would show that the transition continued after the initial surge. Sequential slowing would suggest that annualized summer usage captured a temporary burst, capacity-limited period, or launch cycle.
The composition of that revenue will matter. Continued growth in metered cloud usage would reinforce the platform narrative. A renewed dependence on local deployments would suggest that customized projects remain important for closing large contracts.
The second signal is gross margin. Revenue growth becomes more persuasive if gross profit begins rising at least as quickly as sales.
Improving margin would indicate that utilization, model optimization, and domestic infrastructure are reducing the cost of each paid workload. Flat or declining margin would show that competition and compute requirements are consuming those gains.
Research spending should be evaluated in the same context. Zhipu does not need to stop investing, but revenue must eventually grow faster than the expense base. Otherwise, the platform remains a mechanism for converting capital into subsidized model usage.
The third signal is customer behavior after the next major competitor release. DeepSeek, Qwen, Kimi, and MiniMax have each shown that a single model launch can attract developers quickly.
Zhipu will need to demonstrate that existing production customers stay on GLM services even when rivals publish stronger benchmarks or lower prices. Stable token usage after those launches would support the view that reliability, tooling, and integration create meaningful retention.
Abrupt usage changes would weaken it. They would show that model APIs remain interchangeable and that revenue follows short product cycles rather than lasting platform relationships.
These signals matter beyond China. Zhipu is testing whether a frontier-model company can build a large hosted business using lower-cost access, open models, and domestic computing infrastructure.
Its first-half report provides credible evidence of demand. It does not settle the economic argument.
The 400% increase marks a real business-model change, not merely a publicity milestone. Yet Zhipu’s losses, research bill, and crowded competitive field show how expensive that change remains.
Developers and enterprise buyers should watch what happens after the headline fades. Does recognized API revenue catch up with the announced run rate? Does gross margin improve as token volume grows? Do customers remain after the next model cycle?
Those answers will reveal whether Zhipu has built a recurring cloud platform or entered a faster version of the same capital-intensive race. The company has established momentum. Its next results must show that momentum can survive competition and finance the infrastructure beneath it.



