Alibaba Says Qwen Matches Claude, Raising the Stakes for Anthropic Google
Alibaba says Qwen3.8-Max can rival Claude after completing one autonomous research task over five days, escalating the Anthropic Google model race. The claim puts a Chinese system beside leading American models on long-running work, not only short benchmark questions. Yet Alibaba’s own tests cannot settle whether Qwen performs as reliably under normal business constraints.
The announcement matters because Qwen is becoming more than Alibaba’s answer to a chatbot. The company is developing models, agent software, cloud services, and chips as one connected stack. Qwen3.8-Max is designed to plan, use tools, process feedback, and continue working through complex assignments with limited supervision.
Claude remains the clearest target. Anthropic has made coding and sustained computer work central to Claude’s identity, while Google is connecting Gemini to a vast cloud and productivity footprint. Alibaba is now arguing that capability at this level no longer belongs exclusively to American laboratories.
That argument carries strategic weight, but benchmark parity is not operational parity. Different time limits, token budgets, tools, and scoring rules can produce dramatically different results. The real test is whether Qwen can finish useful work with predictable cost, speed, and failure rates.
What Alibaba Actually Released
Qwen3.8-Max shifts Alibaba’s pitch from answering difficult questions to completing extended assignments.
Alibaba introduced Qwen3.8 as a model family built for coding, professional work, research, and long-horizon agent tasks. A long-horizon task requires the model to maintain a plan across many actions, tool calls, observations, and corrections.
The official Qwen3.8 repository says the model improves autonomous planning and its handling of environmental feedback. It also supports adjustable reasoning effort, allowing developers to control how much computation the model applies before responding.
That positioning is important. Most familiar AI evaluations present a contained prompt and expect one answer. An agent receives a broader objective, chooses intermediate actions, inspects results, and revises its approach when something fails.
According to Alibaba, one internal experiment required Qwen3.8-Max to reconstruct the experiments from a mathematics research paper. The model reportedly worked for about 125 hours without human intervention, then proposed improvements to the original work.
Another showcased task involved chip design. These examples are meant to demonstrate endurance and coordination, not only factual recall. They also resemble the complex assignments that Anthropic has used to frame Claude as a system for serious knowledge work.
Alibaba describes Qwen3.8-Max as its strongest model in the family. The company also released Qwen3.8-2.4T-A95B, an open-weight mixture-of-experts model with 2.4 trillion total parameters and 95 billion activated for each token.
A mixture-of-experts model routes each input through a selected portion of the network. That design can provide very large total capacity without activating every parameter during every step.
The open release separates Qwen’s strategy from a purely closed API approach. Developers can inspect, adapt, and operate released weights under their applicable license terms. Qwen3.8-Max itself is also available through Alibaba services and compatible API formats.
Alibaba’s earlier Qwen3-Max release already showed the direction. Its official model overview described a model with more than one trillion parameters, trained on 36 trillion tokens. It emphasized coding, tool use, long context, and agent performance.
Qwen3.8 pushes that strategy further. Alibaba is making a specific claim about sustained execution, the area where premium models increasingly compete for enterprise demand.
The headline comparison with Claude therefore rests on more than a single leaderboard position. Alibaba wants developers to believe Qwen can receive a substantial project and keep advancing without constant human rescue.
That is the exact change behind the news. A Chinese model provider is no longer arguing only that its system answers the same questions. It is arguing that the system can perform the same category of work.
Why Long-Running Agents Change the Competition
The contest now concerns dependable task completion, which places direct pressure on Anthropic and Google’s enterprise strategies.
Anthropic has built much of Claude’s recent momentum around software development and computer-based work. Claude can inspect repositories, edit files, run commands, analyze errors, and continue through multi-step coding tasks when connected to an agent harness.
Google approaches the same market with Gemini models, Google Cloud, Workspace, and developer infrastructure. Its advantage is distribution across tools that companies already use, from documents and email to data platforms and cloud services.
Alibaba brings a different combination. It controls a major Chinese cloud platform, develops Qwen models, supplies workplace software, and invests in its own AI infrastructure. That gives it several routes for turning model capability into recurring use.
This is why the Anthropic Google comparison matters more than an isolated benchmark victory. All three companies want their models to become the reasoning layer inside software development, research, office workflows, and business operations.
A model that works for minutes can assist with a task. A model that operates reliably for hours or days can take responsibility for a meaningful portion of a project. That difference changes what companies might automate.
Consider a developer asking an agent to update a mature application. The work can require reading hundreds of files, locating dependencies, changing code, running tests, tracing failures, and preparing documentation. Success depends on preserving context across the entire sequence.
Research work creates similar demands. An agent might need to locate papers, reproduce calculations, compare methods, document conflicting evidence, and revise a report. A fluent first answer has little value if the process quietly goes off course after several hours.
These workflows reward persistence, but persistence also multiplies risk. A mistaken assumption made near the beginning can contaminate dozens of later steps. Longer execution creates more opportunities for security problems, wasted computation, and convincing but incorrect output.
That tension gives Alibaba’s announcement significance. If Qwen delivers Claude-like completion rates under real constraints, buyers gain another credible supplier for agent systems. If it only reaches impressive results through unusually generous test budgets, the comparison weakens.
The competitive pressure is immediate for Anthropic because Alibaba names Claude as the capability reference. Google also faces pressure because Gemini competes for the same enterprise workflows and cloud spending.
Open-weight availability adds another dimension. Some organizations want to run models within their own infrastructure, customize behavior, or retain more control over sensitive data. A Qwen model that approaches closed-model performance can make that option more practical.
Developers still must compare deployment requirements, governance controls, security, and support. Open weights do not automatically produce an easier or safer system. They do, however, expand the range of technical and commercial choices.
For knowledge workers, the important development is not a new chat interface. It is the possibility that several model families can manage extended bodies of material and produce auditable deliverables.
That makes source organization increasingly important. A knowledge blending workflow can keep internal material accessible while teams evaluate outputs from different models. The model can change, while the organization’s working context remains under its control.
The Anthropic Google Lead Now Faces Alibaba’s Full Stack
Alibaba is challenging the American leaders with a coordinated model, cloud, software, and semiconductor strategy.
The primary contest is Alibaba versus the Anthropic Google model lead in long-running agent work. Claude and Gemini remain separate products from separate companies, but together they define much of the American reference point Alibaba wants to challenge.
Anthropic brings a focused model business, strong developer adoption, and a reputation centered on demanding reasoning and coding tasks. Google brings research depth, global infrastructure, consumer distribution, and established enterprise relationships.
Alibaba’s response is vertical integration. The company can connect Qwen to Alibaba Cloud, workplace products, consumer services, and domestically developed computing infrastructure. Each layer can reinforce the others.
This full-stack approach became clearer before Qwen3.8 arrived. In May, Alibaba announced Qwen3.7-Max alongside new cloud systems and its Zhenwu M890 AI processor.
Alibaba said an internal test had Qwen3.7-Max work for 35 consecutive hours, make more than 1,000 tool calls, and develop a computing kernel. The company claimed that output exceeded the chip manufacturer’s version, though the result was not independently verified.
The same agent infrastructure announcement described a server containing 128 AI accelerators. Alibaba also said its Zhenwu M890 chip carries 144 GB of on-chip memory and offers three times its predecessor’s performance.
These figures show where Alibaba believes the market is going. Agents generate irregular, sustained demand because they repeatedly call models, retrieve information, run tools, and inspect outputs. Serving those workloads requires more than training a capable model.
A vertically integrated supplier can optimize hardware, inference software, model behavior, and cloud scheduling together. It can also use model demand to support cloud growth, while cloud customers provide workloads that improve agent products.
Anthropic follows a more partnership-oriented infrastructure model. Google already controls its own accelerators and cloud platform, making it structurally closer to Alibaba. The key difference is that Alibaba can optimize for China’s market and hardware constraints.
American export controls have limited Chinese access to some advanced chips. That pressure has encouraged Chinese companies to improve model efficiency, diversify hardware, and develop local semiconductor alternatives.
Alibaba’s model releases therefore carry two messages. The first concerns capability: Qwen can compete with leading systems. The second concerns resilience: Alibaba can deliver an increasingly complete AI stack without depending on every American technology layer.
This does not mean the stacks are equivalent. Anthropic and Google benefit from extensive international developer communities, global cloud regions, mature enterprise programs, and integrations across widely adopted software.
Alibaba also faces the difficulty of earning trust outside its home market. Corporate buyers examine data residency, compliance, support, procurement risk, and geopolitical exposure alongside benchmark performance.
Within China, however, Alibaba’s distribution can turn technical progress into usage quickly. Cloud customers, merchants, developers, and workplace users already operate within parts of its business network.
The result is a two-level contest. Model laboratories compete over coding and reasoning scores, while technology groups compete over who can deploy those abilities at scale.
Alibaba does not need to defeat every Claude or Gemini configuration to change the market. It needs to become credible enough that developers test Qwen, enterprises negotiate among suppliers, and competitors answer its release cadence.
That is already a form of pressure. The frontier becomes harder to define when several systems can solve similar tasks under different operational conditions.
What the Benchmark Headlines Do Not Show
Alibaba’s claim remains provisional because model rankings can reverse when evaluators change time, token, and tool limits.
A benchmark is a standardized test intended to make systems comparable. Agent benchmarks are harder to interpret because the model’s result depends on its surrounding harness, permitted tools, reasoning budget, and execution time.
Alibaba’s launch materials reportedly gave some coding tests a five-hour timeout. PaperBench, which measures attempts to reproduce AI research, allowed some runs to continue for as long as 12 hours.
An independent evaluation discussed in a benchmark analysis used a wall-clock limit between 45 and 60 minutes. Under that tighter constraint, Qwen3.8-Max reportedly placed in the middle of the tested group at its strongest setting.
Both outcomes can reflect real behavior. A model given more time can inspect additional files, attempt more fixes, and recover from earlier mistakes. That does not make the result invalid, but it changes the question being answered.
A five-hour test asks whether the model can eventually solve the assignment with a broad execution allowance. A one-hour test asks whether it can produce value within a typical operational deadline.
Token budgets create the same problem. Reasoning models can generate large amounts of hidden or intermediate computation before returning a final answer. Higher budgets can improve accuracy, but they also increase latency and resource consumption.
A model can therefore look efficient by its public access terms yet become expensive during repeated, reasoning-heavy attempts. The meaningful measure is often total cost per accepted result, including failed runs and human review.
Benchmark design also affects results. A coding model might perform well when given a familiar repository structure, a specific tool set, or an evaluator aligned with its training. Performance can fall when the environment changes.
Vendor tables add another uncertainty. A company controls the tested model version, prompt format, reasoning setting, sampling parameters, and sometimes the comparison configuration. Small methodological choices can alter close results.
Alibaba’s claim should therefore be read precisely. The company says Qwen3.8-Max reaches Claude-like capability on selected tests and extended demonstrations. It has not established universal parity across coding, research, security, speed, or production reliability.
Independent tests can reduce this uncertainty, but they do not eliminate it. Public leaderboards compress many behaviors into a single score, while real deployments depend on specific tasks and failure tolerance.
A customer-support agent must follow policy and avoid unauthorized actions. A coding agent must preserve tests and protect credentials. A research agent must distinguish evidence from inference and retain usable citations.
The best model for one workflow can be a poor choice for another. Companies should build evaluations from representative tasks, then measure completion quality, elapsed time, resource use, and required human intervention.
The 125-hour research example illustrates both promise and concern. Sustained autonomous work sounds impressive, yet five days also provide enormous room for expensive detours. Buyers need to know how frequently the system succeeds and how reviewers detect hidden mistakes.
There is also a safety question. Every extra tool call creates another opportunity for unintended changes or exposure of sensitive information. Long-running agents need permission boundaries, logs, checkpoints, and stopping rules.
These limitations do not erase Alibaba’s achievement. They clarify what remains unproven. Qwen has entered the same category of conversation as Claude, but the outcome now depends on operational evidence.
Why Alibaba Can Sustain the AI Race
Alibaba’s cloud revenue and infrastructure spending show that Qwen is tied to a large commercial strategy, not a temporary research campaign.
Model development at this scale requires continuing investment in chips, data centers, networking, energy, and engineering. Alibaba’s financial results indicate that management is willing to absorb near-term pressure to expand those capabilities.
For the April through June 2026 quarter, Alibaba reported that profit fell 75 percent from the previous year. Revenue increased 9 percent, while revenue from AI cloud and compute services rose 45 percent.
Capital expenditure also increased 75 percent to 67.7 billion yuan. Alibaba said the spending reflected additional computing capacity, expected agent demand, and higher component costs.
The quarterly results provide essential context for Qwen3.8-Max. Alibaba is spending against a theory that AI agents will drive sustained cloud consumption.
That theory makes long-running models commercially attractive. An agent operating across hundreds of steps generates much more inference demand than a chatbot answering one prompt. If customers adopt these workflows, Alibaba can earn from both model access and underlying infrastructure.
The strategy also explains why the company emphasizes complete tasks. Benchmark prestige helps, but finished engineering, research, and office work support a clearer business case for cloud customers.
Alibaba previously committed at least 380 billion yuan to cloud and AI infrastructure over three years. It has also set a goal of exceeding $100 billion in annual AI and cloud revenue within five years.
Those ambitions place Qwen within a broader capital contest. Anthropic depends on major infrastructure partnerships and investor support. Google can fund Gemini through one of the world’s largest advertising and cloud businesses.
The Anthropic Google competition has therefore never been purely technical. Access to compute, customers, distribution, and cash determines how quickly each laboratory can improve models and turn them into products.
Alibaba possesses those assets at considerable scale. Its challenge is converting them into dependable AI services without allowing infrastructure costs to overwhelm returns.
The latest results reveal that tension. AI cloud demand is growing, but higher investment is reducing profit. Management needs Qwen workloads to become durable enough to justify years of spending.
Enterprise adoption will determine whether that calculation works. Large organizations rarely replace a model because of one release-day chart. They test security, data handling, uptime, integration effort, and performance on internal tasks.
Developers will also evaluate the surrounding experience. API compatibility lowers switching friction, but documentation, observability, tool support, and error handling influence whether a model remains in production.
Open models can widen Qwen’s reach beyond Alibaba’s hosted services. The company released its 2.4-trillion-parameter Qwen3.8 model weights in August, followed by a smaller 27-billion-parameter model.
That range allows different adoption paths. Large operators can examine the high-capacity system, while smaller teams can experiment with models that require fewer resources.
An engineering knowledge base can also reduce dependence on any single provider. Teams can preserve documentation and project context, then test several models against the same internal material.
This matters because model leadership changes quickly. A workflow built entirely around one provider’s unique interface can become costly to migrate. A workflow with controlled data and clear evaluations can absorb competitive shifts more easily.
Alibaba’s financial commitment suggests Qwen will remain one of the systems forcing those shifts. Even if its Claude comparison proves too broad, the company has the resources and commercial incentives to keep narrowing the gap.
Three Signals That Will Test Alibaba’s Claude Claim
The next verdict should come from comparable agent tests, production adoption, and direct responses from Anthropic or Google.
The first signal is independent evaluation under matched conditions. Qwen3.8-Max, Claude, and Gemini need the same tasks, tools, time limits, token allowances, and retry policies.
This is the fastest way to test Alibaba’s central assertion. If Qwen remains near Claude across controlled coding and research workloads, the parity claim gains credibility. If its position drops sharply under tighter limits, the launch narrative weakens.
Evaluators should publish more than pass rates. They should report elapsed time, model calls, total tokens, failure categories, and the number of tasks requiring human intervention.
Long-running assignments especially need checkpoint analysis. A model that completes the final task after repeated wrong turns differs from one that follows a stable plan. Both might receive the same pass mark.
The second signal is sustained production adoption. Alibaba needs to show that organizations are using Qwen agents for recurring work, not only controlled demonstrations.
Useful evidence would include completed software changes, research workflows, office processes, and support operations. The strongest cases will disclose review requirements, error rates, and measurable time savings.
Cloud revenue offers another adoption indicator. Continued AI cloud growth would suggest that customers are moving from model trials toward workloads that consume meaningful infrastructure.
Investors should still separate general cloud demand from Qwen-specific adoption. Training projects, storage, and conventional inference can increase revenue without proving that long-horizon agents work reliably.
The third signal is a competitive response. Anthropic can strengthen its position by improving Claude’s task completion, safety controls, or efficiency. Google can answer through Gemini updates and deeper integration with its cloud and productivity products.
A response does not need to mention Alibaba directly. New agent benchmarks, longer autonomous runs, lower failure rates, or broader enterprise controls would show where the competition is moving.
This is where the Anthropic Google pressure becomes visible. When leaders change product priorities after a rival’s release, the rival has influenced the market even before winning a definitive benchmark.
Developers should watch whether agent frameworks make models easier to substitute. Qwen supports common API formats and popular deployment tools, which can reduce the cost of running side-by-side tests.
Enterprises should also watch governance features. Permissions, audit records, execution limits, and human approval gates will matter more as agents remain active for longer periods.
The US-China dimension will remain present, especially around chips, model access, and data rules. Still, treating the story only as a national contest would miss the practical issue facing buyers.
The immediate question is whether another model family can handle consequential work with acceptable reliability. Alibaba says Qwen now belongs in that evaluation set. Independent evidence must determine how often the claim holds.
For developers, the next step is concrete. Select several representative projects, establish identical constraints, and compare accepted outcomes rather than polished demos. Include tasks that require recovery from missing files, faulty assumptions, and tool errors.
For knowledge workers, preserve the sources and decisions behind every extended run. A long answer without a review trail is harder to trust than a shorter result with transparent evidence.
Alibaba has made the Anthropic Google race more crowded and more difficult to summarize with one leaderboard. Qwen3.8-Max now has to convert an ambitious claim into repeatable work.
Will independent tests reproduce Alibaba’s results when time, tools, and reasoning budgets are equal? That answer, not the release headline, will show whether the AI balance has truly shifted.



