top of page

Shanghai Builds a Full-Stack AI Strategy Around Compute, Chips, and Applications

Aug 15
12 min read

Shanghai has turned one Google News headline into a much larger test of whether China can build an AI stack under growing chip constraints. The city is connecting computing capacity, domestic processors, models, and industrial applications instead of betting on one celebrated product.

That strategy took center stage at the 2026 World Artificial Intelligence Conference, held in Shanghai from July 17 to July 20. More than 1,100 companies displayed over 3,000 exhibits, while the program shifted attention from model demonstrations toward infrastructure and deployment.

The tension is straightforward. Shanghai wants to make AI useful across manufacturing, science, transportation, public services, and robotics. Yet the city must support those workloads without dependable access to every advanced processor, fabrication process, or software tool available to American hyperscalers.

This makes the story larger than another local technology initiative. Shanghai is trying to prove that coordinated demand, public computing platforms, alternative chip architectures, and application subsidies can offset disadvantages at the semiconductor frontier.

The Google News Headline Hides a Full-Stack AI Strategy

Shanghai is treating computing power, chips, models, and applications as one connected industrial system.

The immediate event is not a single product release. It is the concentration of several projects, policy measures, and commercial commitments around the city’s AI sector.

At WAIC 2026, Shanghai Unicom announced its UniAI program and said it planned to invest more than 25 billion yuan in intelligent computing infrastructure. The program covers industrial upgrading, public services, urban governance, and other digital applications.

The conference also connected technology suppliers with prospective customers. Its WAIC Connect program organized 172 buyer delegations and presented 63 application scenarios before the event opened.

Those figures matter because they show how Shanghai intends to create demand. The city is not waiting for companies to discover profitable use cases independently. It is identifying buyers, subsidizing experimentation, and placing infrastructure near organizations expected to use it.

The exhibition reflected that same direction. Organizers said the event included more than 300 robots performing tasks related to manufacturing, consumer services, and entertainment. Computing displays expanded beyond individual accelerators to supernodes and large clusters.

A supernode combines many processors through high-speed connections so they can handle large AI workloads as one system. That approach becomes especially important when one domestic accelerator cannot match the performance of the most advanced foreign chip.

The city’s existing base gives the strategy more weight. Shanghai’s AI industry exceeded 450 billion yuan in 2024, according to the city’s AI industry profile. It included more than 380 sizable AI enterprises and over 60 EFLOPS of public intelligent computing capacity.

An EFLOPS represents one quintillion floating-point operations per second. The measure offers a rough view of theoretical capacity, although it does not reveal utilization, networking efficiency, memory performance, or the cost of delivering a useful answer.

Shanghai’s AI sector reportedly exceeded 550 billion yuan in 2025, growing more than 30 percent from the prior year. Local authorities attributed that expansion partly to policies reducing the cost of computing, models, and training data.

This is the important change behind the Google News result. Shanghai has moved beyond presenting AI companies as isolated success stories. It is assembling a system that can move a project from research through computing access, model development, customer testing, and commercial deployment.

The system also stretches beyond central Shanghai. Computing workloads can run in western regions with abundant energy, while developers, chip companies, laboratories, and industrial buyers remain concentrated near the coast.

China’s national computing network is meant to connect those pieces. It links data centers, supercomputing facilities, and edge infrastructure so workloads can move according to latency, energy, and capacity requirements.

That arrangement resembles an electrical grid more than a conventional collection of independent data centers. The analogy is imperfect, but the goal is clear: make processing capacity easier to request, measure, route, and purchase.

Shanghai’s advantage lies in coordination. Its challenge is proving that the coordinated system produces reliable services rather than impressive capacity announcements.

Why Shanghai Is Measuring Demand in Tokens

The city’s infrastructure policy is shifting from theoretical computing performance toward the amount of AI work customers actually consume.

Shanghai highlighted this change at the Intelligent Computing Shanghai Summit on January 27, 2026. Participants described a move from measuring infrastructure mainly through floating-point operations toward tracking token use.

A token is a small unit of text or data processed by a language model. Customers usually consume tokens when they send a prompt, retrieve information, run an agent, or generate an answer.

Token consumption is not a complete measure of economic value. A model can process many tokens while producing little useful work. Still, the metric connects infrastructure more directly with application demand than a processor’s maximum benchmark score does.

The China Academy of Information and Communications Technology joined three major telecommunications operators and local computing providers to launch a citywide collaboration system. The initiative aims to align computing suppliers with model developers and application companies.

This reflects a wider change in AI economics. Training a foundation model remains expensive, but repeated inference now determines whether an AI product can support millions of daily interactions.

Inference is the process of running a trained model to generate an output. It includes everything from answering a short question to operating a software agent through hundreds of intermediate steps.

China’s daily AI token consumption reportedly rose from 100 billion in early 2024 to 140 trillion in March 2026, according to the National Data Bureau. ByteDance’s Doubao alone reportedly increased from about 100 billion daily tokens in May 2024 to more than 120 trillion by March 2026.

Those figures should be treated as reported measurements, not independently audited usage. However, their direction explains why Shanghai is emphasizing infrastructure, network capacity, and operating cost.

Agents consume more tokens than simple chat interfaces because they plan tasks, review intermediate results, call tools, and correct mistakes. A product that appears simple to the user can generate extensive background computation.

This creates pressure on data centers and chip suppliers. Training demand arrives in large, planned bursts, while inference demand can fluctuate throughout the day and spread across thousands of customers.

Shanghai’s public platforms can help smaller developers avoid building their own clusters. Local computing subsidies also reduce the cost of early experiments before a company has stable revenue.

Pudong and Xuhui have offered eligible startups computing, model, and corpus coupons. Each category can provide support of up to 1 million yuan, according to information released by Shanghai authorities in January.

A corpus is an organized collection of text or other data used to train, adapt, or evaluate an AI system. Access to a useful corpus can matter as much as raw processing capacity for specialized industrial applications.

The coupon model does more than lower expenses. It can steer startups toward local infrastructure providers and encourage them to build around supported chips, model services, and datasets.

That creates a strategic feedback loop. Subsidized applications generate demand for domestic computing. Higher demand gives chip and platform companies more opportunities to improve compatibility, performance, and developer support.

The risk is that subsidies can also conceal weak economics. A project may appear viable while credits cover its inference bill, then struggle when it must pay the full operating cost.

Token-based measurement makes that problem easier to see. Officials and investors can compare installed capacity with actual consumption, customer retention, response quality, and revenue-producing tasks.

For enterprise buyers, this is more useful than another headline benchmark. The central question is not how many processors a cluster contains. It is how reliably the system completes work at an acceptable cost.

That is why Shanghai AI computing policy now reaches beyond the data center. It ties infrastructure spending to industrial agents, specialized models, robotics, and software used inside factories.

Domestic Chips Must Compensate at the System Level

Shanghai’s hardest problem is turning varied domestic processors into dependable computing systems that developers can use without constant reengineering.

American export controls have restricted China’s access to some advanced AI accelerators and semiconductor technologies. Those limits have increased pressure on Chinese companies to develop processors, packaging methods, networking systems, and software alternatives.

The comparison with Nvidia remains unavoidable. Nvidia’s advantage is not limited to raw chip performance. Its CUDA software environment, libraries, developer community, and established server designs reduce the work required to deploy large models.

Shanghai cannot replace that environment through one accelerator announcement. It needs a group of chip companies, system builders, cloud providers, model developers, and customers to improve together.

The city already hosts several relevant companies. Biren Technology and MetaX develop general-purpose graphics processors for AI workloads, while Iluvatar CoreX has pursued domestic GPU designs for data centers.

Other Shanghai teams are exploring architectures that avoid a direct contest over the most advanced conventional GPU. Lightelligence develops optical computing systems, which use light for selected data-processing and communication tasks.

The company is jointly deploying a 10,000-GPU supernode cluster in Shanghai. Founder Shen Yichen has described system integration and photonics-electronics packaging as continuing bottlenecks.

That qualification matters. An optical component can perform well in a laboratory while remaining difficult to package, cool, manufacture, program, and connect inside a commercial server.

Dongfang Suanxin has taken another route with the DF1000, a near-memory 3D AI chip shown at WAIC 2026. Near-memory computing places processing elements close to memory to reduce the time and energy spent moving data.

Traditional processors frequently transfer information between separate memory and computing components. That movement creates a bottleneck known as the memory wall, especially during data-intensive AI inference.

The DF1000 stacks components in a three-dimensional structure. The company says the design can improve data movement and reduce dependence on the most advanced manufacturing processes.

That claim has not received broad independent benchmarking. Buyers still need evidence covering application performance, software compatibility, production yield, reliability, and total system cost.

Tongji University has also introduced a specialized Moving Horizon Unit for autonomous vehicles, robots, and drones. The chip is designed to repeatedly observe conditions, calculate a response, and update a machine’s actions.

Such specialized processors illustrate Shanghai’s broader approach. Instead of forcing every workload through one general accelerator, developers can match particular architectures with edge AI, robotics, scientific computing, or industrial control.

Edge AI runs models near the device producing the data, rather than sending every input to a distant cloud. It can reduce latency and network use, although smaller devices usually impose strict limits on memory and power.

The city benefits from having applications close to chip designers. A robotics company can report latency problems, a factory can identify reliability gaps, and a vehicle developer can expose power or thermal constraints.

That feedback can improve a product faster than an isolated benchmark contest. It also lets local suppliers optimize complete systems around available manufacturing technologies.

However, combining more chips does not automatically remove the performance gap. Large clusters need fast interconnects, efficient scheduling, reliable memory, cooling, and software that distributes work without wasting capacity.

A cluster built from less efficient processors can require more electricity and physical space. It can also create more points of failure and impose greater maintenance costs.

Industry estimates cited in recent reporting put Huawei near Nvidia in China’s 2025 AI chip market, with each holding roughly 40 percent. Analysts expected Huawei to gain share as Nvidia’s position weakened under supply restrictions.

Those estimates show progress toward domestic substitution, but they do not prove complete independence. Semiconductor production depends on equipment, materials, packaging, intellectual property, and manufacturing processes distributed across several countries.

Shanghai’s chip push is therefore a system-level bet. The city does not need every domestic component to lead the global market individually. It needs enough components to work together at a cost and reliability level that supports real applications.

Applications Are the Test, Not the Exhibition

Shanghai will strengthen its AI position only if subsidized infrastructure produces repeatable work inside factories, laboratories, vehicles, and public services.

The city’s July policy package for manufacturing offers support for industrial models, AI coding systems, physical AI, agents, industrial software, and specialized datasets.

Qualified core technology projects can receive support of up to 20 million yuan. Security projects for industrial models and agents can receive up to 10 million yuan, while certain dataset platforms can receive up to 20 million yuan.

The measures also encourage industrial computing platforms to offer trial access to low-code agent tools, token credits, and computing support.

These programs target a practical adoption problem. Manufacturers often possess valuable operational data, but that information sits across machines, engineering documents, maintenance records, and internal systems.

An industrial model needs more than public internet text. It must work with domain terminology, equipment states, safety limits, access controls, and the specific workflows used by employees.

A useful factory agent might compare a sensor alert with maintenance history, technical manuals, and recent production changes. It then needs to show its evidence and route the recommendation to an authorized engineer.

That workflow places demands on data quality, retrieval, latency, and accountability. A fluent answer is not enough when an error can halt production or damage equipment.

Organizations considering such systems first need a searchable information layer. A well-maintained engineering knowledge base can help teams organize technical documents before adding automated decisions.

Shanghai has several industries where this approach can be tested, including automotive manufacturing, semiconductors, pharmaceuticals, finance, shipping, and advanced equipment.

Robotics provides another visible use case. More than 300 robots at WAIC performed activities across manufacturing and daily services, but exhibition performance remains different from long-term operation.

A factory robot must repeat tasks across changing lighting, object placement, equipment wear, and human activity. It also needs predictable maintenance and a clear procedure for safe failure.

Scientific applications impose different requirements. Researchers may prioritize data access, reproducibility, specialized numerical methods, and the ability to inspect intermediate results.

Public services introduce privacy and governance concerns. A citywide AI system can affect residents who did not choose the technology and cannot easily switch providers after an error.

These differences explain why one general model will not settle Shanghai’s application challenge. Each sector needs evaluation methods tied to its own risks and desired outcomes.

The most credible indicators will come from sustained use. Buyers should look for lower equipment downtime, faster design cycles, reduced energy consumption, improved inspection accuracy, or shorter research workflows.

Published application counts deserve caution. A registered model, funded pilot, or conference demonstration does not necessarily represent active commercial deployment.

The same concern applies to computing utilization. A city can announce substantial capacity while workloads remain concentrated among a small number of large companies.

Shanghai’s reported 60 EFLOPS of public intelligent computing power shows infrastructure availability. It does not disclose how evenly capacity is used, what customers pay, or how often developers encounter compatibility problems.

Energy creates another constraint. The International Energy Agency expects global data center electricity consumption to rise from about 485 terawatt-hours in 2025 to approximately 945 terawatt-hours in 2030, as described in its data center outlook.

China is placing some computing centers near renewable energy resources in western regions. High-speed network links then connect those facilities with coastal technology hubs.

This can lower pressure on Shanghai’s local grid, but distance affects latency and operational complexity. A large training job can tolerate remote processing more easily than a robot requiring an immediate response.

Developers must decide which work belongs in a western cluster, a Shanghai data center, or an edge device. The answer depends on data sensitivity, response time, energy cost, network reliability, and model size.

The city’s full-stack strategy is designed to support those choices. Its success will depend on whether the connections between layers remain usable after subsidies and conference attention fade.

Three Signals Will Show Whether Shanghai’s Bet Works

The next evidence must come from utilization, commercial chip deployments, and measurable industrial outcomes.

The first signal is sustained token consumption on Shanghai’s public and commercial computing platforms. Rising use would show that models and agents are moving beyond isolated demonstrations.

The strongest version of that signal would include customer retention, workload diversity, and improving cost per completed task. Raw token growth alone can reflect inefficient models or subsidized experimentation.

If utilization remains low despite new capacity, the city may have built infrastructure faster than applications can absorb it. That would weaken the argument that coordinated investment creates its own demand.

The second signal is independent evidence from domestic chip deployments. Shanghai AI chips need public benchmarks covering real models, energy use, cluster scaling, failure rates, and software migration.

Announcements about optical processors, near-memory designs, and domestic GPUs remain meaningful research indicators. Commercial contracts and repeat orders would provide stronger proof.

Watch the 10,000-GPU optical supernode project associated with Lightelligence. A stable deployment would support Shanghai’s argument that alternative architectures can improve large computing systems.

Watch the DF1000 as well. Customer testing could reveal whether its near-memory design delivers practical benefits outside controlled demonstrations.

The third signal is measurable value from industrial applications. Shanghai’s manufacturing policy should produce pilots involving agents, datasets, robotics, industrial software, and physical AI.

The city needs to disclose outcomes that buyers can evaluate. Useful measures include production uptime, inspection performance, inference cost, energy consumption, employee adoption, and time saved.

A long list of funded projects would not settle the issue. The stronger result would be multiple organizations renewing deployments with their own budgets.

Competition will shape all three signals. Huawei, Alibaba, Baidu, and other large Chinese technology companies can supply models, clouds, chips, and developer platforms outside Shanghai’s local programs.

American companies remain relevant even when their latest hardware is restricted. Their software practices, model releases, and infrastructure economics continue setting benchmarks for global buyers.

Shanghai must therefore compete on more than political support or installed capacity. It needs predictable tools, credible security, accessible data, affordable inference, and systems that engineers can maintain.

The city enters that contest with meaningful advantages. It has a large industrial economy, research institutions, chip designers, model companies, investors, and government buyers in close proximity.

It also faces real weaknesses. Domestic processor supply remains uneven, alternative architectures need commercial validation, and subsidized pilots can obscure customer willingness to pay.

That combination makes Shanghai’s experiment worth following. It is testing whether a regional industrial cluster can compensate for constrained access to frontier hardware through coordination across the entire AI stack.

The next Google News headline will probably highlight another chip, robot, cluster, or funding program. Readers should look past the individual announcement and ask three questions.

Is the new capacity being used after trial credits expire? Are domestic processors completing real workloads with acceptable reliability and energy use? Are industrial customers expanding deployments with their own money?

Those answers will determine whether Shanghai has built a durable AI advantage or an ambitious collection of infrastructure projects. Track the deployments, compare the operating data, and treat every showcase claim as the start of verification rather than its conclusion.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page