top of page

Anthropic Classified AI Workloads Are Leaving the Pentagon Despite Its Court Win

5 hours ago
14 min read

Anthropic classified AI workloads are almost gone from Pentagon systems, despite the company winning a major court ruling only two weeks earlier. The Department of Defense has transferred roughly 90% of that work to other model providers, Pentagon technology chief Emil Michael said on September 10. Officials expect the remaining migration to finish before October.

That timing creates an unusual split between legal victory and operational defeat. A federal judge ruled in August that one government effort to label Anthropic a supply-chain risk was unlawful. Yet the Pentagon has continued removing Claude from classified environments, including systems where Anthropic once held a rare early advantage.

The dispute is no longer only about one contract or one AI safety policy. It has become a test of who controls the rules after a commercial model enters a military network. Anthropic wants restrictions covering mass domestic surveillance and fully autonomous weapons. Pentagon leaders insist that military systems must support all lawful uses under government authority.

OpenAI, Google, xAI, and other providers now have the opportunity to absorb the work Anthropic is losing. The transition suggests that technical quality alone does not secure a lasting place inside classified infrastructure. Contract terms, deployment control, and policy alignment can matter just as much.

The Pentagon Has Already Moved 90% of the Work

The migration has advanced far beyond planning, making Anthropic’s removal an operational reality rather than a negotiating threat.

Michael, the under secretary of defense for research and engineering, said approximately 90% of the relevant classified work had moved to alternative providers. He also said Anthropic technology had been removed from the Maven Smart System months earlier.

Maven combines data, software, and artificial intelligence to support military analysis and operational decisions. Claude entered those workflows through Anthropic’s partnership with Palantir, which develops and operates key parts of the system.

The transition marks a sharp change from early 2026. At that point, Claude was widely reported as the only frontier commercial model operating across some Pentagon classified environments. Frontier models are highly capable, general-purpose AI systems trained to perform many language, analysis, and coding tasks.

Anthropic had reached that position through years of security and deployment work. Its models operated through protected infrastructure rather than the public Claude service used by consumers. Classified deployment also required government accreditation, controlled access, and isolation from ordinary commercial networks.

The relationship expanded in July 2025, when the department awarded Anthropic a prototype agreement with a ceiling of $200 million. The two-year arrangement covered AI applications in intelligence, operations, cybersecurity, and other national security functions, according to the company’s defense agreement.

OpenAI, Google, and xAI received comparable agreements for frontier AI projects. Each agreement carried a ceiling of $200 million, although a ceiling is not guaranteed spending. An official contract notice described work across warfighting and enterprise domains.

Those parallel awards gave the Pentagon a route away from dependence on one model provider. However, signing prototype agreements did not instantly make every model ready for classified use. Each provider still needed secure infrastructure, technical integration, testing, and permission to process information at the required classification level.

In May, the department expanded its classified AI relationships to eight companies while excluding Anthropic from the new group. The participants included major model developers and infrastructure providers. That expansion created a broader replacement pool for Claude and reduced the operational risk of moving workloads.

The latest 90% figure indicates that these alternatives progressed quickly. It does not reveal how much classified work exists, how the Pentagon measures a workload, or which provider received each migrated task. A workload might represent an application, a model endpoint, a mission process, or another unit defined internally.

That missing denominator matters. Moving 90% of many small tasks differs from transferring the most sensitive or technically difficult applications. The department has not publicly released a detailed inventory, which is understandable for classified systems but limits outside assessment.

Still, the direction is unambiguous. The Pentagon has largely dismantled Anthropic’s early position inside its classified AI stack. The final portion is expected to follow before October, even as the legal dispute surrounding the government’s treatment of Anthropic continues.

Why Anthropic Classified AI Workloads Became a Policy Fight

The central conflict concerns authority: Anthropic wants narrow contractual limits, while the Pentagon rejects a private vendor’s continuing control over lawful military uses.

The relationship deteriorated over two categories of use. Anthropic opposed using Claude for mass surveillance of Americans and for fully autonomous weapons that select and engage targets without meaningful human involvement.

CEO Dario Amodei said in February that the company would not remove those safeguards. He argued that some uses could undermine democratic values, while others exceed what current AI systems can perform safely and reliably.

Anthropic did not broadly reject defense work. The company supported intelligence analysis, cybersecurity, logistics, operational planning, and other military applications. Its objection focused on a limited set of high-consequence uses where model errors or unrestricted data processing could produce irreversible harm.

Pentagon leaders saw the issue differently. They argued that elected officials and military commanders, not technology companies, must determine how lawful operations are conducted. From that perspective, an external vendor should not retain an effective veto after its model enters an accredited government environment.

The disagreement also involved technical control. Anthropic told courts that it could not secretly alter a model after deployment inside an isolated classified system. Once the government received and hosted an approved version, the company said, Anthropic lacked the remote access needed to disable it during an operation.

That claim challenged one rationale behind the supply-chain risk designation. Supply-chain controls normally address risks such as hidden access, sabotage, compromised components, or foreign influence. Anthropic argued that a contractual disagreement about acceptable use did not establish any such technical threat.

The Pentagon maintained that reliability and vendor behavior can also affect national security. Officials questioned whether a supplier that might object to future operations could remain dependable across urgent military scenarios. They framed unrestricted lawful use as a readiness requirement rather than permission for unlawful conduct.

Both positions contain a real governance concern. A private company should not quietly determine national defense policy. At the same time, a government contract does not automatically erase negotiated safety limits, constitutional safeguards, or the technical limitations of probabilistic models.

The phrase “all lawful uses” does not resolve that tension. Legality can depend on facts, jurisdiction, executive authority, and interpretations that change over time. A model vendor also cannot verify the legality of every classified application it never sees.

Technical reliability creates another complication. Large language models generate outputs from statistical patterns and can produce confident errors. In administrative work, a human can often catch and correct those mistakes. In targeting, surveillance, or weapons control, the same failure can carry much greater consequences.

Anthropic’s recent military evaluations underscore that concern. The company reported that advanced models can assist with tactical intelligence targeting and conventional weapons development. It also found capabilities that could help malicious actors identify targets or improve weapon performance.

Those findings do not establish that Claude was used unlawfully inside the Pentagon. They demonstrate why both parties treated usage conditions as more than abstract ethics language. As model capabilities improve, the gap between analytical assistance and operational decision-making becomes harder to define.

The Pentagon chose to resolve the immediate control question through replacement. By shifting classified work to vendors that accepted its terms, the department reduced its reliance on a supplier whose policies might conflict with future missions.

That choice also imposed a clear market signal. Companies seeking classified defense work must now assume that policy alignment will be evaluated alongside model performance, security accreditation, and price. Safety commitments that remain under vendor control can become grounds for exclusion.

OpenAI and Other Providers Gain More Than New Contracts

Anthropic’s exit gives competing labs a chance to become infrastructure providers for missions that are difficult to move once deeply integrated.

The immediate beneficiaries include OpenAI, Google, xAI, and other companies added to the Pentagon’s classified AI program. Each brings different models, cloud relationships, deployment methods, and acceptable-use policies.

OpenAI became the clearest comparison after announcing a classified deployment agreement during the initial dispute. The company also identified limits involving domestic mass surveillance and autonomous weapons, but its negotiated language satisfied Pentagon officials.

That outcome shows how small differences in contract wording can produce large commercial consequences. Anthropic and OpenAI publicly described overlapping concerns. Yet one company lost its classified workloads while the other gained entry to the same environment.

The distinction appears to involve who interprets and enforces the limits. Anthropic sought restrictions that it considered clear and durable. OpenAI reached terms that left lawful-use decisions within the government’s chain of command while preserving stated safety principles.

Outside observers cannot fully compare the agreements because classified implementation details remain unavailable. It is therefore too simple to conclude that one provider abandoned safeguards or that another retained operational control. The practical difference lies in what the Pentagon accepted.

Google and xAI further reduce the department’s dependence on any single laboratory. Multiple model providers let agencies match workloads to particular capabilities, compare outputs, and switch systems when performance changes.

Michael has argued that model rankings move frequently and that leading systems should become more comparable over time. If that judgment proves correct, the Pentagon gains leverage. It can treat frontier models as replaceable components instead of unique strategic dependencies.

Classified integration is not completely interchangeable, however. A model must connect with secure data, applications, identity systems, and human workflows. Teams also develop prompts, evaluation sets, review procedures, and software around a provider’s behavior.

Replacing a model can change output quality even when benchmark scores appear similar. It can also require new testing for hallucinations, security weaknesses, bias, latency, and performance on specialized military language.

The migration from Claude therefore demonstrates significant integration capacity. Moving 90% of workloads in several months suggests that the Pentagon and its contractors had either limited dependency or strong incentives to rebuild quickly.

It also reflects a deliberate multi-vendor strategy. The department’s classified expansion brought additional model and infrastructure companies into protected environments after the Anthropic conflict began.

For competing labs, access creates more than short-term contract value. Classified work exposes models to demanding analytical settings and specialized user needs. Feedback from those deployments can influence evaluation methods, product development, and future government procurement.

The opportunity also carries risk. Any provider can face the same underlying conflict if its safety policy later diverges from a military application. Signing an agreement does not eliminate uncertainty about how more autonomous models will be used.

Enterprise buyers should notice this dynamic. AI systems are becoming embedded in workflows where replacement costs extend beyond an API change. Organizations must account for provider policy, contract stability, deployment control, and model portability before concentrating critical work with one vendor.

That lesson applies even outside defense. A regulated company might lose access because of compliance rules, litigation, export controls, or a provider policy change. Teams need an accurate map of which workflows depend on which models.

Maintaining that map requires more than a procurement spreadsheet. Technical decisions, model evaluations, meeting notes, and policy approvals often sit across separate systems. A searchable knowledge base can preserve the reasoning needed when a model must be replaced quickly.

For Anthropic, the competitive cost is substantial even without public spending totals. Claude went from being the leading commercial model inside classified Pentagon workflows to holding a shrinking share. Rivals now get the deployment experience, institutional relationships, and integration opportunities that Anthropic helped establish.

A Court Win Did Not Restore Anthropic’s Pentagon Position

The August ruling limited the government’s punitive authority, but it did not require the Pentagon to keep buying or deploying Claude.

U.S. District Judge Rita Lin ruled that one effort to designate Anthropic as a supply-chain risk was unlawful. She found that the government had retaliated against the company for criticizing its approach to military AI use.

The ruling rejected the idea that officials could invoke national security without an adequate factual connection to sabotage or technical compromise. According to the court, the record did not establish that Anthropic intended to interfere with its models inside government systems.

The decision gave Anthropic an important legal and reputational victory. A supply-chain risk label can affect far more than one agency contract. It can discourage contractors and commercial customers from working with a designated company, even when a restriction formally applies only to certain government business.

More than 100 enterprise customers reportedly contacted Anthropic with concerns after the government’s actions. The company argued that the measures threatened billions in potential revenue and damaged its reputation beyond defense procurement.

An August federal ruling blocked key parts of that campaign. Anthropic said afterward that it wanted to work productively with the government while supporting national security.

However, procurement discretion is different from punitive designation. A court can stop an agency from applying an unlawful label without ordering that agency to select a particular vendor. Anthropic has said its lawsuits do not seek to force the government to contract with it.

That distinction explains how the Pentagon can continue the migration. Officials can choose other providers based on contract compatibility, operational preference, redundancy, or mission requirements. The government does not need to restore Claude merely because one legal mechanism was invalidated.

A separate legal proceeding also remains relevant. Anthropic challenged another Pentagon action under a different statutory framework in Washington. The two cases overlap in subject but do not necessarily produce identical remedies.

The administration itself has sent mixed signals. Commerce Secretary Howard Lutnick said in early September that the government trusted Anthropic. Michael responded that the company remained a supply-chain risk for the Defense Department and its industrial base.

That disagreement reveals institutional complexity rather than a settled government-wide position. One agency can view Anthropic as a valuable American AI company while another rejects it for a specific mission relationship.

The Pentagon’s continued migration also reduces the practical value of any later Anthropic victory. Restoring eligibility does not automatically restore technical integrations, user adoption, or contractor momentum after replacement systems take over.

This creates the article’s central reversal. Anthropic successfully defended its right not to be branded a national security threat without sufficient evidence. Yet it has not preserved the classified business that triggered the fight.

The outcome can encourage AI companies to challenge government retaliation while still warning them about the commercial cost of disagreement. Legal protections can constrain punishment, but they cannot guarantee demand.

There is also a risk of overstating what the 90% migration proves. It does not demonstrate that replacement models perform better than Claude. It does not show that Anthropic’s safeguards were unnecessary. It only shows that the Pentagon found alternatives it considered acceptable enough to deploy.

Likewise, the court ruling does not validate every Anthropic policy choice. It addresses the lawfulness of government action and the supporting record. It does not decide how much control an AI provider should retain over classified military applications.

Those unresolved questions will return as models become more capable. Future systems may plan complex operations, coordinate software agents, analyze surveillance feeds, or recommend actions with limited human prompting. Contractual language written for chatbots may not adequately govern those uses.

The Pentagon has addressed its immediate dependency through diversification. The larger dispute over responsibility remains open.

The Migration Trades Vendor Control for Integration Risk

Removing Anthropic resolves one policy conflict, but it creates new testing, oversight, and accountability burdens across a larger provider network.

Classified migration is not equivalent to switching a public chatbot subscription. Each model operates within a controlled technical environment and touches data that cannot move freely between systems.

Before a replacement handles live work, teams must test whether it follows instructions, cites evidence, protects sensitive context, and resists manipulation. They must also determine when a human reviewer needs to intervene.

Model substitution creates behavioral differences. Two systems can receive the same prompt and produce different conclusions, uncertainty levels, or refusals. Those variations matter when analysts use AI to summarize intelligence, identify patterns, write code, or prioritize information.

Evaluation is especially difficult because classified data cannot be sent to public benchmark services. The Pentagon must build internal test sets and keep them current as models and threats change.

A multi-provider environment expands that burden. Each model version can have separate failure modes, policy controls, context limits, and infrastructure requirements. Updating one component may require renewed security and performance testing.

The department gains resilience by avoiding dependence on Anthropic, but it also assumes greater integration complexity. Vendor diversity helps only when systems remain observable and workloads can move without losing essential context.

The most sensitive question concerns operational accountability. If an AI output contributes to a military decision, officials need to know which model produced it, which data informed it, and which human approved the resulting action.

That chain becomes harder to reconstruct when several providers and contractors participate. Logging, provenance, access control, and version tracking must follow the workload across every transition.

The government must also distinguish assistance from delegation. A model that searches documents or drafts a briefing supports human judgment. A model that selects targets, coordinates autonomous systems, or ranks individuals for surveillance moves closer to exercising consequential authority.

Anthropic drew its red lines around that boundary. The Pentagon argued that existing law, policy, and command responsibility already governed such decisions. Neither position removes the need for technical controls that make human oversight measurable.

Another risk is policy drift. A provider might accept current terms but revise its safety framework after a new model reveals unexpected capabilities. The government could also expand an application beyond the use originally reviewed.

Contracts need change-control procedures for those cases. Otherwise, the same disagreement can recur after a different vendor becomes deeply embedded.

The Pentagon’s approach also affects the broader defense industrial base. Contractors build applications on commercial models, sometimes through cloud platforms or intermediary software. A provider dispute can therefore propagate through several layers of a program.

Clear inventories and portability plans can reduce that exposure. Teams should retain model-independent datasets, evaluation criteria, and workflow documentation where security rules permit. They should also avoid tying essential business logic to undocumented quirks of one model.

This does not mean every organization needs several live providers. Multi-model architecture adds cost and complexity. The lesson is to know what replacement would require before a policy or legal shock makes it urgent.

Anthropic’s experience offers a particularly sharp example. Its technical lead in classified deployment did not protect it when acceptable-use terms became incompatible with the customer’s requirements. The Pentagon’s successful migration suggests that even an established frontier model can be displaced under enough institutional pressure.

However, speed should not be confused with validation. Public reporting does not reveal comparative error rates, security findings, user satisfaction, or mission outcomes after the transition. Those gaps prevent a confident judgment about whether the replacement improved operational performance.

The final 10% might also be the hardest portion. Residual workloads often remain because they depend on specialized integrations, unique model behavior, or unresolved accreditation. A deadline can drive completion, but it cannot guarantee equivalent capability.

That is why October should be treated as a migration milestone, not proof that the underlying tradeoffs have disappeared.

What to Watch Before and After October

Three signals will show whether the Pentagon completed a durable transition or merely replaced one unresolved governance problem with several new ones.

The first signal is the department’s confirmation that all Anthropic classified AI workloads have moved. Officials should clarify whether “all” includes every Claude-dependent application, contractor integration, and classified endpoint.

A complete migration would reinforce the Pentagon’s claim that frontier models are increasingly substitutable. A delay would suggest that Anthropic retained technical advantages or deeper dependencies than the 90% figure reveals.

The most useful evidence would go beyond a percentage. The department could disclose, without exposing classified missions, how it defines workloads and whether replacement systems passed comparable evaluations.

The second signal is the next legal decision involving Anthropic’s remaining challenge. The August ruling constrained one government action, but litigation over a separate designation continued.

Another Anthropic victory would weaken the Pentagon’s attempt to characterize a policy dispute as a supply-chain threat. A government victory would preserve broader authority to exclude vendors whose conduct officials consider a national security risk.

Neither result would automatically restore Claude to classified systems. The legal outcome would instead shape the leverage available in future negotiations between agencies and AI developers.

The third signal is whether competing providers adopt or revise military-use limits. OpenAI’s agreement already showed that a company can state safety boundaries while accepting the Pentagon’s demand for governmental authority.

Future policies from OpenAI, Google, xAI, and other participants will reveal whether that balance survives real deployments. A new dispute would show that Anthropic’s conflict reflected a structural problem rather than one company’s negotiating position.

Model capability will raise the stakes. Anthropic’s evaluations indicate that advanced systems are becoming more useful for targeting analysis and weapons development. As those capabilities spread, vague contract language will become harder to defend.

Enterprise technology leaders should follow these signals even if they never handle classified data. The Pentagon is stress-testing questions that commercial buyers will also face: who controls a deployed model, how quickly it can be replaced, and what happens when vendor policy changes.

Those questions belong in procurement reviews before an AI tool becomes critical infrastructure. Buyers should document approved uses, technical dependencies, evaluation results, and exit procedures. They should also identify who can authorize a change when performance, law, or provider policy shifts.

The Anthropic classified AI workloads migration delivers a clear near-term result. The Pentagon appears ready to remove Claude from its protected networks before October, while competitors take over the work.

The harder question remains unsettled. Should an AI company preserve narrow control over high-risk uses, or should a lawful government customer hold complete authority after deployment?

Watch the final migration report, the remaining litigation, and the policies adopted by Anthropic’s replacements. Together, they will show whether this episode established a stable model for military AI or only postponed the next confrontation.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page