People's Daily Backs a Chinese Name for Token, but Usage Will Decide the Winner
- Martin Chen
- 22 hours ago
- 12 min read
People's Daily has turned a four-month-old token translation decision into a fresh dispute about technological influence, despite developers already having an established vocabulary.
The immediate development is an August 6 commentary arguing that choosing a Chinese term for token affects who defines the language of artificial intelligence. The piece appeared on The Paper's hot list, although its aggregator supplied no verified publication time.
The underlying terminology decision is not new. China's national authority for scientific terminology published the Chinese name in March 2026. The current debate concerns whether that official term will gain practical authority beyond government documents and domestic media.
That distinction matters. Technical language becomes influential when researchers, developers, vendors, educators, and customers use it consistently. An official decision can accelerate adoption, but it cannot make a disputed definition technically precise.
The contest is therefore not simply English against Chinese. It is institutional terminology against working language, with engineers and markets deciding whether the two eventually converge.
The Token Decision Happened in March, Not August
The hot-list story revives an existing terminology campaign rather than announcing a new technical standard.
On March 25, China's National Committee for Terms in Sciences and Technologies published a trial terminology notice. It selected a Chinese expression meaning roughly "lexical unit" for token in the AI field.
The committee is authorized to review and publish standardized scientific terminology. It was established in 1985 with State Council approval and works under a system jointly supported by Chinese science institutions.
The word "trial" is important. The committee's stated process allows a newly released term to gather feedback before formal approval. The March notice was therefore an institutional starting point, not proof of universal adoption.
Two days earlier, National Data Administration director Liu Liehong used the same expression at the China Development Forum. He described a token as both a measurable unit and a connection between technical supply and commercial demand.
A March 25 government summary then presented the Chinese name as settled. It also linked the terminology to claims about the scale of domestic AI activity.
According to the administration, daily token calls in China rose from 100 billion in early 2024 to 100 trillion by the end of 2025. It said the figure exceeded 140 trillion in March 2026.
Those numbers imply growth of more than one thousand times in roughly two years. They also explain why officials view the unit as more than specialized computer science vocabulary.
However, the metric needs careful handling. "Token calls" is not a standard international measure with a single public audit method. Different organizations can count model input, generated output, cached content, internal reasoning, or agent activity differently.
The figures show how Chinese authorities frame AI consumption. They do not independently establish model quality, economic value, or comparative international market share.
The August commentary adds a political and cultural argument to that March decision. Its central claim is that naming affects technological discourse, public understanding, and eventually rule-making power.
That makes the commentary newsworthy, but it does not change how models split text. No tokenizer changed because an editorial endorsed a new term.
The practical event is a coordinated effort to move a Chinese technical expression from formal documents into mainstream public language. The tension begins when that effort meets established developer habits.
Why Token Became a Strategic Word
Token now sits where language, computing costs, product limits, and AI infrastructure meet.
A token is a unit that a model processes after a tokenizer converts text or other input into numerical pieces. It can represent a word, part of a word, punctuation, code, or another learned unit.
That definition is deliberately broader than "word." The English word "unbelievable," for example, can become several pieces under one tokenizer and fewer under another.
Chinese text creates another complication. A token can represent one character, multiple characters, punctuation, a byte sequence, or part of a less common expression.
Tokenization is the mechanism that performs this conversion. Modern systems often build vocabularies from frequently recurring byte or character patterns instead of relying on dictionary words.
The influential subword research behind byte-pair encoding showed why such pieces help language systems handle rare and previously unseen words. Later model families adapted related methods at far greater scale.
The mechanism gives token a precise technical role. It also gives the unit several commercial roles that did not exist when tokenization mainly interested language-processing researchers.
Model providers use tokens to describe context limits. They use tokens to meter API consumption. Developers track tokens to control latency, capacity, and spending.
AI agents increase that significance because one visible request can trigger many hidden operations. An agent can retrieve documents, call tools, revise a plan, and produce several intermediate model responses.
Each step can consume input and output tokens. The final answer might contain only a few paragraphs, while the supporting process handles a much larger volume.
That creates a genuine reason for policymakers to seek accessible vocabulary. People cannot evaluate AI packages or infrastructure claims if the basic unit remains opaque.
China's official campaign also arrived during a shift from model demonstrations toward operational AI. The National Data Administration described tokens as a settlement unit connecting technical capacity with commercial demand.
That framing contains a useful insight. A token can function like a metering unit inside a model service, even though it is not a currency or a universal measure of intelligence.
The distinction is essential. One token from one model does not necessarily equal one token from another model in informational content, processing effort, or output quality.
Token counts depend on the tokenizer. Costs depend on model architecture, hardware, caching, batch size, service design, and the balance between input and output.
A million tokens processed by a small classification model cannot be compared directly with a million tokens processed by a large reasoning system. The visible quantity is similar, but the work differs.
This is where official language can influence policy. Once governments use a term in procurement, statistics, education, and infrastructure planning, that term shapes what institutions measure.
The risk is that linguistic clarity can produce false economic precision. A standardized name does not standardize the underlying tokenizers, workloads, or accounting methods.
The People's Daily argument is strongest when it asks who makes technical concepts accessible. It becomes weaker if terminology is treated as evidence of technical leadership by itself.
Official Language Meets Developer Language
The primary contest is between institutional adoption and the vocabulary that developers already use every day.
Most Chinese AI developers recognize the English word token. It appears in API documentation, software libraries, research papers, dashboards, error messages, and international technical discussions.
That installed base gives the English term considerable staying power. Developers do not choose vocabulary only for cultural reasons. They also optimize for interoperability and searchability.
A programmer debugging a context-limit problem will often search the exact phrase shown by an API. If the interface says "input tokens," translating the query can reduce the quality of available results.
Research creates similar pressure. Papers written in English use token across natural language processing, computer vision, speech systems, and multimodal models.
The word also travels with related terms such as tokenizer, token embedding, token budget, token probability, and special token. Replacing one element does not automatically produce a coherent technical family.
However, leaving everything in English creates a different barrier. Policymakers, business buyers, students, and ordinary users can mistake token for a cryptocurrency asset or authentication credential.
Computer security already uses token for credentials and access mechanisms. Blockchain markets use the same word for digital assets. AI uses it for model input and output units.
A clear Chinese term can reduce that ambiguity inside a specific context. It can also make teaching and public policy less dependent on unexplained English vocabulary.
The official expression follows a familiar Chinese construction. Its first element refers to a linguistic unit, while the second suggests a basic component.
That makes the term readable, but not perfectly complete. AI systems tokenize more than words, and modern multimodal systems process image, audio, and video representations alongside text.
Even within text, a token does not always align with a lexical unit. Byte-level systems can divide an unfamiliar character or symbol into pieces that carry no independent lexical meaning.
This technical mismatch explains why some specialists proposed alternatives. The disagreement is not automatically resistance to Chinese terminology. It can reflect competing views of what the unit represents.
One approach emphasizes language. Another emphasizes symbols or model-level units. A third simply keeps the English word because its meaning changes across technical domains.
Official institutions face a difficult design problem. A term must be accurate enough for experts, understandable enough for the public, and flexible enough to survive technical change.
Developers face a more immediate test. The chosen word must help them communicate without creating extra translation work between Chinese documentation and global systems.
Both sides have legitimate needs. Institutional language favors consistency and broad access. Working language favors compatibility, speed, and precision within an existing technical network.
The likely outcome is not total replacement. Chinese publications can pair the local term with token, while code, parameters, and international research retain English identifiers.
That pattern already appears in official coverage. The National Data Administration and other state outlets commonly display both expressions together instead of deleting the English term.
Bilingual usage weakens the idea of a winner-takes-all contest. It also offers a practical route for the Chinese expression to gain recognition without isolating domestic developers.
Token Naming Does Not Create Technical Power
A country gains technological influence when its terminology describes systems that others use, measure, and build upon.
The People's Daily commentary connects naming with discourse power. History gives that argument some support, but it also sets a demanding standard.
Terms shape public understanding. They determine what appears in textbooks, regulations, procurement documents, media coverage, and executive discussions.
A country that participates early in technical naming can prevent important concepts from entering public debate through inconsistent translations. It can also reduce dependence on explanations created elsewhere.
Yet vocabulary follows technical institutions as often as it leads them. Terms spread internationally when they travel with widely used products, influential papers, standards, and open-source projects.
China's strongest claim to AI influence therefore comes from models, research, applications, chips, data infrastructure, and developer communities. A translation can organize that activity, but it cannot substitute for it.
Consider how developers encounter tokenization. They experience it through model behavior, context limits, API meters, multilingual performance, and unexpected truncation.
If a Chinese model processes Chinese text with fewer units while preserving output quality, developers will care. The associated terminology can then acquire concrete technical meaning.
If Chinese standards bodies publish measurement rules that make token statistics comparable across vendors, enterprise buyers will care. That would give the vocabulary an operational function.
If Chinese researchers influence multilingual tokenizer design, the country can shape how global models represent languages. That work carries more technical weight than choosing a local label alone.
Language representation is not neutral. A tokenizer trained around some languages can split other languages into longer sequences, increasing context usage and processing requirements.
The exact effect varies by model and vocabulary. It should be measured on comparable tasks rather than asserted from a language's writing system.
This creates a credible policy agenda behind the terminology debate. Chinese institutions can fund multilingual benchmarks, require clearer token accounting, and support research on efficient representation.
They can also push vendors to disclose which tokenizer they use, what counts toward usage, and how cached or internal tokens appear in reports.
Those actions would make the official term part of a measurable governance framework. Without them, the campaign risks becoming a branding exercise around an unstable unit.
The reported 140 trillion daily calls illustrate the problem. The number sounds exact, but readers cannot compare it confidently without scope and methodology.
Does it include consumer chat products, enterprise APIs, internal batch processing, and synthetic-data generation? Does it count repeated agent steps or only external calls?
Do all reporting companies define an input token the same way? Are cached tokens included? How are multimodal units converted into the aggregate?
Public documentation reviewed for this article does not answer those questions. The gap does not make the statistic false, but it limits what analysts can conclude.
A defensible claim is that domestic model activity increased sharply according to Chinese government data. A stronger claim about global leadership requires comparable definitions and independent evidence.
The same standard should govern terminology. Calling a unit strategic does not make every measurement based on it strategically meaningful.
What the Token Debate Still Gets Wrong
The argument over naming compresses several unresolved technical and economic questions into one emotionally appealing symbol.
The first problem is definitional scope. A token in a text model is not always a word-like unit, while a token in an image model can represent a patch or learned visual element.
Speech and video models introduce further variations. Multimodal architectures can transform several input types into internal sequences whose units do not share one intuitive public meaning.
A translation centered on words works well when explaining chatbots to new users. It becomes less exact when applied across the entire AI field.
The second problem is comparability. Token counts differ across tokenizers, and providers can expose different portions of the inference process.
Reasoning models complicate the picture further. A provider can bill or report internal reasoning units separately, include them in output, or hide part of the process behind product abstractions.
Agent systems add tool descriptions, retrieved passages, conversation history, and intermediate results. A short user request can therefore create a large context repeatedly.
This matters for enterprise buyers. They need to compare completed tasks, accuracy, latency, and total resource use, not only raw token volume.
A model that uses fewer tokens but produces unreliable work is not automatically more efficient. Another model might consume more units because its tokenizer divides the same content differently.
The third problem is political overreach. Critics can reasonably accept Chinese technical terminology while rejecting the claim that one translation materially changes international influence.
The August discussion drew skeptical online reactions from people who viewed the campaign as disconnected from working developer practice. Social media is not representative evidence, but it identifies the adoption challenge.
Users repeatedly compared token with familiar abbreviations such as CPU or USB. Those examples show that local-language explanations and English technical labels can coexist for decades.
The fourth problem is the temptation to treat usage volume as economic output. Tokens are inputs and outputs of computation, not revenue, productivity, or verified customer value.
Automated systems can generate enormous token volumes through testing, synthetic data, repeated agent loops, or inefficient prompts. High usage can indicate adoption, waste, or both.
A useful metric would connect token volume with completed work. Examples include resolved service requests, accepted code changes, reviewed documents, or tasks completed without human correction.
Knowledge workers also experience tokens indirectly. A long personal knowledge workflow can consume context through retrieval and repeated analysis without exposing the underlying unit.
Clear vocabulary helps those users ask better questions about limits and data handling. Still, they should not need to become token accountants to evaluate whether an AI product works.
The fifth problem is durability. Model architectures can reduce the importance of today's token boundaries or move more processing into continuous internal representations.
Tokens remain central to current language models, but the public term must survive changes in architecture. A narrow definition may age quickly.
The committee's trial process provides room to address this concern. Feedback from linguists, AI researchers, educators, developers, and vendors can test the chosen expression across domains.
The decision should be judged by whether it improves communication without concealing technical variation. Cultural symbolism is relevant, but precision remains the harder requirement.
Three Signals Will Show Whether the New Term Matters
Adoption, measurement standards, and international technical output will reveal whether the naming campaign produces lasting influence.
The first signal is consistent product adoption. Watch Chinese model vendors, cloud providers, telecommunications companies, universities, and developer platforms.
A term has moved beyond policy when independent organizations use it without needing an official prompt. Documentation matters more than ceremonial speeches because developers consult it during real work.
Interfaces offer an even stronger test. If dashboards, billing records, usage alerts, and procurement contracts adopt the Chinese term, it becomes part of operational AI.
Bilingual labels would still count as adoption. They can preserve international compatibility while making the unit legible to a broader domestic audience.
If major vendors continue using only English in developer-facing systems, the official expression will remain strongest in government and media language. That would weaken claims of broad technical authority.
The second signal is a shared accounting standard. China needs published rules explaining how organizations calculate aggregate token activity.
Those rules should separate input, output, cached content, internal reasoning, agent loops, and multimodal units. They should also explain how statistics combine providers with different tokenizers.
Comparable measurement would strengthen the March usage claims. It would help enterprise buyers evaluate efficiency and help policymakers distinguish valuable adoption from automated volume.
Without such rules, a larger number can reflect a different counting method. The national metric will attract attention while offering limited analytical value.
The third signal is international uptake of Chinese technical work. Watch tokenizer research, multilingual benchmarks, model documentation, standards proposals, and open-source libraries.
A Chinese name does not need to replace token in English to matter internationally. It can support domestic education while Chinese researchers influence the underlying mechanisms.
The strongest evidence would be external adoption of Chinese-developed methods for multilingual token efficiency, model metering, or cross-vendor reporting.
That would connect language with technical contribution. It would also turn discourse power from an editorial claim into an observable result.
Over the next three months, the terminology committee's feedback process deserves close attention. Any clarification of scope would show whether officials recognize the multimodal and accounting problems.
Vendor documentation should follow. Consistent use across unrelated companies would strengthen the case that the term has escaped a policy-driven media cycle.
Measurement disclosure is the final test. If agencies publish a methodology for the 140 trillion figure, analysts can evaluate what the growth actually represents.
People's Daily is right that technical language has consequences. Names influence who can participate in a discussion and how institutions organize a new market.
But vocabulary earns authority through use. The Chinese name for token will matter when it improves education, documentation, measurement, and model design.
Developers and enterprise buyers should therefore watch behavior instead of rhetoric. Does the term clarify a bill, a context limit, or a procurement requirement?
Does it help compare competing systems? Does it support better multilingual models? Those outcomes will determine whether the campaign expands technological influence or merely renames an imported unit.