U.S. Data Vendors Face Scrutiny Over Chinese AI and Pentagon Work
Google News surfaced a sharp conflict involving four U.S. data companies reportedly serving Chinese AI labs while pursuing or holding American government work.
The companies are Surge AI, Mercor, AfterQuery, and Turing. A Forbes investigation linked them to customers including Tencent, Ant Group, Alibaba, and ByteDance. Some of those relationships were documented through internal communications, project records, and interviews with industry sources.
The reporting does not establish that the vendors transferred classified information, military data, or proprietary datasets belonging to American model developers. It raises a different question. Should companies that build specialized training data for U.S. agencies also supply expertise to organizations Washington considers security risks?
Rep. Michael McCaul, a Texas Republican and longtime China hawk, argues that government access should come with restrictions. He told the New York Post that vendors should not help a Pentagon-designated Chinese military company improve its AI while seeking U.S. contracts.
That position turns an obscure supply-chain issue into a procurement debate. Washington has spent years restricting advanced chips, manufacturing equipment, and investment. Specialized training data still moves across borders with far less scrutiny.
The central conflict is therefore not simply America versus China. It is the commercial promise of a global data market versus the security obligations attached to government work.
What the Data Vendors Reportedly Sold
The controversy concerns packaged human expertise, not a reported transfer of Pentagon files or secret American model weights.
Modern AI systems need more than raw text gathered from websites. Developers also purchase carefully designed examples, expert evaluations, grading instructions, and corrections that teach models how to perform difficult tasks.
This process is often called post-training. It improves a model after its broad initial training by showing it stronger answers, identifying errors, and rewarding desired behavior.
Data vendors recruit software engineers, scientists, accountants, lawyers, and other specialists for that work. They then organize the experts’ output into datasets that AI labs can use repeatedly.
The recent investigation describes a market in both custom projects and off-the-shelf datasets. Custom data is created for a particular customer or model objective. Off-the-shelf data is standardized and can be sold to several buyers.
That distinction matters because reusable datasets can spread more than isolated facts. They may include task designs, evaluation methods, quality controls, and reasoning patterns developed through earlier projects.
According to the underlying training data investigation, Chinese buyers sought material covering areas such as finance, cybersecurity, coding, and model self-improvement. The publication reviewed communications involving buyers and workers at several data platforms.
Forbes reported that executives at Tencent described relying on Surge AI for data. It also linked Mercor to Tencent, AfterQuery to Ant Group and Alibaba, and Turing to ByteDance.
Those findings require careful attribution. The publication said it reviewed internal messages, transcripts, and project documentation, but most underlying records were not published in full. Readers cannot independently inspect every document supporting each commercial link.
The companies also offered limited public responses. Mercor declined to comment, while AfterQuery and Surge said they do not disclose customer information. Turing said it serves both proprietary and open-source model development.
The Chinese companies named in the report did not provide responses for that investigation. Their silence neither confirms nor disproves the reported relationships.
The scale estimates also come from industry sources rather than public financial filings. Two data entrepreneurs reportedly placed annual spending by six leading Chinese AI labs at about $500 million with American data-labeling companies.
The same reporting attributed at least $50 million in recurring Chinese revenue to AfterQuery. It said Chinese labs represented 2% of Mercor’s second-quarter revenue, citing a person familiar with the matter.
These figures show why the trade attracts vendors, if accurate. The market is large enough to influence sales strategy, yet private enough that customers and projects remain difficult to track.
The four vendors are not accused in the reporting of violating an existing export rule. Specialized datasets generally do not face the same controls applied to advanced processors or semiconductor equipment.
That legal gap is central to the story. A transaction can be commercially routine under current rules while still attracting scrutiny from lawmakers responsible for national security policy.
The public evidence supports a narrow conclusion: American vendors reportedly supply structured AI training materials to Chinese commercial organizations. It does not support claims that they handed over U.S. military intelligence or copied specific customer datasets.
That boundary should remain clear. Otherwise, a legitimate debate about procurement and export controls can become an unsupported allegation of espionage.
Why Google News Is Amplifying a Pentagon Procurement Problem
The political pressure comes from the vendors’ overlapping customer lists, not from data-labeling work alone.
Surge AI has a documented history of small Defense Department awards. A federal Small Business Innovation Research record lists a 2023 Army project for an AI-assisted data-labeling platform.
That Army data award had a total value of $149,285. The public record identifies Surge Labs as the awardee and Edwin Chen as the principal investigator.
A separate Air Force project awarded Surge $73,194 in 2023. Its stated objective involved a language model for searching and summarizing intelligence reports.
These awards are small beside major defense technology contracts. However, they establish that Surge entered the defense procurement system for work involving data infrastructure and intelligence-related applications.
The public award pages do not show that Surge received classified data. They also do not prove that any method, material, or employee crossed between its American and Chinese projects.
McCaul’s argument instead focuses on institutional alignment. A vendor trusted with government work, he contends, should not simultaneously strengthen an organization designated by the Pentagon.
Tencent is the most important example. The U.S. government has previously identified Tencent as connected to China’s military, a characterization the company has disputed.
Designation does not automatically make every commercial transaction with Tencent illegal. It does, however, create reputational and policy risks for an American contractor serving defense customers.
The same issue now surrounds Alibaba. In June 2026, the Pentagon placed Alibaba, Baidu, BYD, and other businesses on an expanded list of entities it considers Chinese military companies.
The official military company list says Alibaba contributes to China’s defense industrial base through government affiliations. Alibaba rejects that conclusion and has challenged its designation.
The list itself is a government determination, not a criminal judgment. Inclusion can block defense contracting and bring further restrictions, but it does not establish that every product from a listed company serves military purposes.
That nuance matters for AfterQuery’s reported relationship with Ant Group. Ant operates independently from Alibaba in important corporate respects, although their businesses and histories remain connected. A reported project involving Ant should not automatically be described as direct work for China’s military.
The Google News headline nevertheless captures the political problem effectively. Washington is asking technology suppliers to choose whether government work carries obligations beyond minimum legal compliance.
Federal procurement already imposes security rules, conflict disclosures, and supply-chain requirements on many contractors. Congress could expand those rules to cover AI training data or work for designated foreign entities.
Such a change would pressure vendors before regulators define a new export category. A company might retain the legal ability to sell a dataset abroad while losing eligibility for federal contracts at home.
That approach would make procurement policy a substitute for export controls. It could move quickly because agencies already possess broad authority to evaluate contractor risk.
It would also create complications. AI data vendors use distributed workforces, subcontractors, and expert marketplaces. Determining where a dataset originated, which customer funded its design, and whether its methods were reused can be difficult.
Government buyers would need more than a customer blacklist. They would need clear rules for segregating teams, data, evaluation methods, and reusable materials.
Without that detail, restrictions could become inconsistent. One agency might reject a vendor based on a disputed customer relationship, while another continues buying from it.
The immediate pressure therefore falls on Pentagon procurement officials. They must decide whether a contractor’s foreign commercial work presents an operational risk, a political risk, or neither.
The Real Export Gap Is Human Expertise
Chip restrictions target computing capacity, while the reported data trade transfers methods for making that capacity more useful.
The United States has concentrated its China technology policy on processors, semiconductor tools, and manufacturing knowledge. Those controls reflect the enormous computing demands of training advanced AI systems.
Compute is only one input. Model developers also need data that targets weaknesses, measures performance, and teaches reliable behavior in specialized domains.
Raw internet text cannot fully prepare a model to complete an advanced accounting analysis or review laboratory procedures. Those tasks require examples and judgments from people who understand the work.
A vendor might recruit a professional to produce a solution, another expert to critique it, and a third to score competing answers. The resulting package teaches a model what strong performance looks like.
Evaluation rubrics are especially valuable. A rubric defines the qualities an answer must contain and the errors that should lower its score. It turns professional judgment into a repeatable training signal.
The reported trade suggests Chinese labs can purchase this infrastructure instead of building every expert network and quality process internally. That can reduce experimentation and shorten development cycles.
Nathan Lambert, formerly a research scientist at the Allen Institute for AI, told Forbes that high-quality data becomes especially important as models move into harder domains. His observation explains why datasets command significant commercial value.
Yet the national-security case is not automatic. Data vendors often create material from public knowledge and paid expert labor. Restricting those services involves different legal and practical questions from restricting a specialized chip.
A processor is a physical product with a model number, shipment record, and measurable performance. A training dataset can be copied, modified, combined with other material, or delivered through a service contract.
The same vendor may also produce distinct data for different customers. Serving Anthropic on one project does not prove that an off-the-shelf dataset sold elsewhere contains Anthropic’s confidential information.
This distinction is essential. The controversy involves shared production capacity and expertise, according to the reporting. It does not establish the unauthorized resale of an American lab’s proprietary customer data.
There is still a strategic concern. A vendor learns which expert workflows produce useful improvements, how to reject weak contributions, and how to build evaluation systems efficiently.
Even if no confidential record changes hands, that operating knowledge can help another lab spend its resources more effectively. The advantage lies partly in the production method.
Washington has encountered similar challenges in other technology sectors. Export controls work best when officials can define the restricted item and identify its destination.
Knowledge-intensive services are harder to isolate. Regulators must decide whether to control the final dataset, the subject matter, the production process, the buyer, or some combination.
Overly broad rules could also damage American vendors. Global sales support revenue, recruit international experts, and help U.S. companies maintain influence over industry standards.
Sean Cai, an AI data consultant quoted by Forbes, warned that sweeping restrictions could weaken open-source development and reinforce the position of a few closed model providers. He did not dismiss security concerns, but questioned simple restrictions.
That argument defines the strongest counterposition to McCaul’s proposal. A forced commercial split could limit Chinese access, but it could also fragment research markets and reduce competition among American providers.
Open-source models add another layer. Their weights and training methods can circulate broadly, while the datasets used to improve them may remain private.
Restricting American data services would not eliminate Chinese alternatives. It would encourage domestic data companies, universities, and crowdsourcing platforms to replace U.S. suppliers.
The policy question is therefore about delay and leverage, not permanent denial. Restrictions might slow access to certain expert networks or evaluation methods. They cannot erase the underlying knowledge.
Readers should also separate data quality from model capability. Purchasing strong examples can improve performance, but it does not guarantee a competitive frontier model.
Labs still need computing infrastructure, researchers, algorithms, deployment systems, and sustained testing. Data is a high-value input, not a complete substitute for those capabilities.
The Google News attention matters because it moves this subtle mechanism into public debate. Readers familiar with chip controls can now see why human-generated training material belongs in the same strategic conversation.
The Evidence Does Not Yet Prove a Security Breach
The strongest verified concern is an undisclosed policy conflict, while the most serious allegations remain unproven.
No public evidence currently shows that Surge AI, Mercor, AfterQuery, or Turing transferred classified U.S. information to a Chinese customer.
The reporting also does not demonstrate that a dataset created under a Pentagon contract was resold abroad. It does not show that Anthropic, Google, Meta, or OpenAI customer materials were shared without authorization.
Those would be substantially more serious claims. They require contracts, technical records, or direct testimony that has not been made public.
The available evidence instead combines reported customer relationships, government award records, industry estimates, and comments from a lawmaker.
That package supports scrutiny. It does not justify describing the companies as conduits for military secrets.
The distinction could become difficult to maintain as political attention rises. Headlines compress complicated supply chains into a simple image of companies “feeding” two opposing sides.
That framing communicates the conflict, but it can imply a common pool of information moving between customers. Data vendors say little publicly about how they separate projects, making the implication hard to test.
The companies could address much of the uncertainty through transparency. They could disclose policies for designated entities, customer screening, dataset provenance, and separation between government and commercial teams.
They could also explain whether off-the-shelf products incorporate techniques developed during confidential projects. Such explanations would not require revealing customer data.
Independent audits would provide stronger evidence. An auditor could test access controls, employee permissions, data lineage, and contractual restrictions without publishing sensitive materials.
Government procurement officers should ask similar questions. A foreign customer relationship is only one part of a risk assessment.
They should examine whether personnel overlap, whether tools share storage environments, and whether reusable datasets preserve confidential specifications. They should also verify how contractors manage subcontractors.
The vendors’ limited responses leave those questions open. Confidentiality obligations can prevent them from naming customers, but they can still publish general safeguards.
The Chinese side also disputes Washington’s broader designations. Alibaba and Baidu have denied that they are military companies, while China’s government accuses the United States of stretching national-security definitions.
The Associated Press reported that the Pentagon’s list expanded to 188 entities in 2026. The designation dispute illustrates why policymakers should distinguish official classification from independently proven military activity.
The issue is not merely semantic. A rule tied to the Pentagon list would inherit disputes, additions, removals, and court challenges associated with that list.
A vendor might lose federal eligibility when a customer is designated, even if the underlying commercial project remains lawful and unrelated to defense.
Congress would need to decide whether restrictions apply retroactively. It would also need a process for companies to disclose, wind down, or challenge problematic relationships.
Another uncertainty concerns Mercor’s reported federal contract. Public reporting says the company told employees it had secured government work, but available details remain limited.
Without an agency name, scope, award record, or contract terms, it is premature to place Mercor’s situation on the same evidentiary footing as Surge’s published awards.
AfterQuery and Turing present different cases again. The reports connect them to Chinese customers, but the public evidence described so far does not establish Pentagon contracts for either company.
Treating all four vendors as identical would obscure meaningful differences. Each relationship requires separate verification.
McCaul’s proposal is a political position, not an existing government-wide rule. Agencies have not publicly announced a blanket ban covering every data vendor that serves a designated Chinese company.
Congress has, however, intensified its attention to Chinese AI systems and American technology dependencies. A July 2026 House investigation examined how U.S. companies use Chinese open-weight models.
That investigation concerns the opposite direction of travel, Chinese models entering American systems. Together, the two debates show lawmakers examining AI supply chains in both directions.
One side asks whether American firms should buy Chinese models. The other asks whether American firms should sell Chinese labs the data that improves those models.
Neither question can be settled through nationality alone. Risk depends on the buyer, application, dataset, access controls, and sensitivity of the underlying expertise.
The strongest policy would focus on those measurable factors. A broad political label is easier to communicate, but less precise to enforce.
Three Signals Will Show What Happens Next
The next phase will be defined by procurement rules, company disclosures, and evidence about dataset separation.
The first signal is formal action from the Pentagon or Congress. McCaul has described a principle, but agencies must translate that principle into contract language before it changes vendor behavior.
A meaningful action would identify covered entities and restricted services. It would also explain whether a vendor can preserve eligibility by ending work, separating operations, or adopting stronger controls.
Rules limited to Pentagon-designated companies would be narrower than a prohibition covering all Chinese AI customers. That distinction would determine the commercial impact.
A government-wide procurement clause would strengthen the argument that Washington sees training data as a strategic asset. Continued case-by-case treatment would indicate that officials remain uncertain about workable boundaries.
The second signal is disclosure from the named vendors. Surge AI, Mercor, AfterQuery, and Turing can clarify their positions without exposing confidential customer records.
The most useful disclosures would describe screening standards, prohibited applications, data segregation, and reuse policies for off-the-shelf products.
If the companies announce that they avoid designated military-linked customers, McCaul’s pressure will have produced a market response without legislation.
If they defend global sales and emphasize existing controls, the debate will shift toward whether those controls satisfy government buyers.
Silence carries its own cost. Federal customers usually demand confidence about supply-chain risk, especially when contractors work with intelligence, defense, or sensitive operational information.
The third signal is independent evidence about data provenance. Provenance records track where training material came from, who modified it, and which restrictions govern its use.
A credible audit showing strict project separation would weaken suggestions that American customer knowledge flows directly to Chinese labs. It would not resolve the broader dispute over selling expertise abroad.
Evidence of reused confidential specifications would produce a much more serious controversy. It could trigger contract enforcement, litigation, or export-control action beyond procurement policy.
Readers should watch for documents, not just stronger rhetoric. Contract awards, agency memoranda, audit reports, and published compliance policies will reveal more than anonymous market estimates.
The same discipline applies when following the story through Google News. Aggregated headlines help identify an emerging dispute, but they often collapse allegations, verified records, and political proposals into one sentence.
The verified record is narrower. Surge received two small defense research awards in 2023. Forbes reported commercial relationships between four U.S. data vendors and several Chinese technology organizations.
The Pentagon has designated Tencent and several other Chinese technology companies as military-linked entities, although the companies dispute those descriptions. McCaul wants federal contractors to choose between U.S. government work and certain Chinese customers.
What remains unknown is just as important. The public record does not show classified transfers, cross-customer dataset reuse, or a current blanket prohibition covering these sales.
For developers, the story illustrates why data provenance is becoming a product requirement. Teams buying training material need to understand its origin, licensing, reuse conditions, and quality controls.
Enterprise buyers face a similar responsibility. A vendor’s customer list can become a compliance issue even when the purchased service performs well.
Knowledge workers who create training examples should also ask how their contributions can be packaged and resold. Their work may travel across models, markets, and jurisdictions in ways that standard contractor agreements do not make obvious.
The conflict will not disappear if Washington restricts four companies. Demand for expert-generated data is expanding because models need better feedback in specialized fields.
Policy can influence which suppliers capture that demand and how transparently they operate. It cannot remove the economic value of professional judgment.
The most useful next step is therefore careful scrutiny rather than immediate certainty. Follow the procurement language, examine the vendors’ safeguards, and separate documented transactions from claims about national-security consequences.
Will Washington define AI training data as a controlled strategic resource, or continue using individual contracts to police the market? The answer will determine whether this Google News controversy becomes a lasting policy shift or another unresolved warning about AI’s hidden supply chain.



