top of page

Anthropic Mythos EU Access Arrives Three Months Late

2 hours ago
12 min read

Anthropic gave the European Union access to Mythos 5 after more than three months of negotiations, despite previously signaling that access was coming. The Anthropic Mythos EU access agreement lets the European Union Agency for Cybersecurity, known as ENISA, test the restricted cybersecurity model. It does not include the newer Mythos 5.1.

That distinction turns a delayed handover into something more consequential. ENISA can now examine an important frontier model, but its evaluation already trails Anthropic’s release cycle. The company introduced Mythos 5 in June and followed it with Mythos 5.1 on September 1.

The gap places Anthropic’s safety commitments against the practical limits of independent oversight. European regulators want direct evidence about systems that can find and potentially exploit software vulnerabilities. Anthropic controls when those regulators receive the models, what version they receive, and which safeguards remain active.

The dispute is therefore not simply about whether Europe obtained access. It concerns whether public evaluators can examine frontier systems while their findings still apply to the models being deployed. OpenAI has already given ENISA access to GPT-5.6 Cyber and GPT-6 Astra, according to European officials.

Anthropic Mythos EU Access Covers the Older Model

ENISA has started testing Mythos 5, but Anthropic has already moved its most sensitive model line to Mythos 5.1.

European Commission spokesperson Thomas Regnier confirmed the new arrangement on September 10. He said ENISA had received Mythos 5 and was testing it following discussions with Anthropic. The ENISA access arrived more than three months after talks began.

Mythos 5 is a restricted version of the same underlying model used for Claude Fable 5. Anthropic removed or reduced certain cybersecurity safeguards for vetted organizations. Those changes let approved defenders conduct work that the generally available model would block.

That work can include vulnerability research, exploit validation, and security testing on authorized systems. These activities have legitimate defensive value, but the same abilities can support attacks when misused. Anthropic therefore distributes Mythos through controlled programs instead of offering unrestricted public access.

The company launched Mythos 5 on June 9 as an upgrade for existing Project Glasswing participants. Project Glasswing is Anthropic’s partnership program for organizations protecting critical software and infrastructure. Anthropic said it planned to add partners periodically and build a broader application process.

The Mythos 5 launch described the model as limited to Glasswing partners and selected researchers. It also said expansion would happen in consultation with the United States government. That condition placed access decisions inside a mixture of corporate review, national security policy, and technical risk management.

ENISA was expected to become part of that expansion. However, expectation did not produce immediate access. Brussels kept pressing the company while European officials questioned whether restrictions treated a close United States partner unfairly.

The delay also crossed a major product boundary. Anthropic released Mythos 5.1 on September 1, nine days before ENISA’s access became public. Anthropic says the newer model improves cybersecurity and biology performance, while remaining available only to vetted organizations.

ENISA therefore received a relevant model rather than the newest one. Testing Mythos 5 can still reveal how Anthropic’s access controls, safeguards, and cyber capabilities behave. Yet those conclusions cannot automatically be transferred to Mythos 5.1.

Even small model updates can alter agent behavior, refusal patterns, tool use, and vulnerability discovery. A regulator needs the evaluated version to match the deployed version closely. Otherwise, oversight becomes a historical review instead of a current risk assessment.

The Anthropic Mythos EU access decision resolves the immediate question of whether ENISA can test something. It does not resolve whether ENISA will receive timely access to each significant successor.

The Three-Month Delay Is the Central Safety Problem

Independent evaluation loses value when access arrives after the model generation has already changed.

Frontier AI testing depends on timing because capabilities can move faster than formal policy processes. A model evaluator needs enough time to design tests, reproduce failures, and communicate findings before broad deployment decisions become difficult to reverse.

ENISA now faces that problem with Mythos 5. Its work can measure the June model’s ability to discover vulnerabilities, propose exploits, or navigate complex systems. It can also examine whether Anthropic’s monitoring and access restrictions contain those abilities.

However, European officials must interpret the results against a September product landscape. Mythos 5.1 is already Anthropic’s leading Mythos-class system. Anthropic says that version contains gains in cybersecurity and biology, two areas with significant dual-use concerns.

Dual-use capabilities serve both legitimate and harmful purposes. A vulnerability discovery system can help a maintainer patch software before attackers act. The same system can help an attacker locate weak code and create a working exploit.

Anthropic addresses that conflict through different access layers. Fable models use stronger safeguards for general availability. Mythos models relax selected safeguards for approved organizations with defensive or scientific purposes.

That structure makes vetting a central safety control. Access is not a minor administrative detail added after model development. It determines which organizations can observe the model’s least restricted behavior and whether independent institutions can challenge Anthropic’s claims.

The delay shows that a trusted-access program can create its own oversight bottleneck. Anthropic must assess organizations, satisfy government requirements, and prevent dangerous access. Regulators must negotiate from outside that process, even when their mandate involves evaluating the resulting risks.

European institutions also possess enforcement authority under the EU AI Act. That role can complicate voluntary information sharing because evidence gathered during testing might inform later regulatory action. A developer could reasonably protect sensitive systems while still creating an appearance of selective oversight.

Neither concern removes the timing problem. If companies provide older models after newer versions launch, public testing bodies remain one step behind. The system can appear cooperative while limiting the practical reach of that cooperation.

This matters beyond one model. Frontier developers increasingly release systems in rapid sequences, with capability updates arriving between formal evaluations. A three-month negotiation can consume an entire product generation.

The problem becomes sharper for cyber models because defenders and attackers operate continuously. A delayed evaluation cannot recover the defensive time lost before access. It also cannot establish whether the latest model introduced new behaviors.

Anthropic’s own strategy recognizes the importance of a defender advantage. Project Glasswing was designed to give selected security organizations an early opportunity to find and repair vulnerabilities. That advantage depends on who receives access and when.

ENISA protects European networks, supports incident coordination, and helps develop cybersecurity policy. Excluding it from early testing weakens the geographic reach of the defender advantage. Granting access later improves coverage, but it does not recreate the missing early window.

The central question is not what Anthropic Mythos is in abstract terms. The relevant question is whether access governance can keep pace with a model designed for urgent security work.

Anthropic’s Safety Promise Meets Europe’s Oversight Mandate

The main conflict pits Anthropic’s controlled-access safety model against Europe’s demand for independent and current evidence.

Anthropic argues that direct access creates greater misuse risk than narrowly delivering defensive outputs. A malicious operator with open-ended prompting can probe a model, refine requests, and pursue harmful objectives. A user receiving only a patch or alert has less freedom.

The company has started placing Mythos behind specialized security products. Claude Security can scan customer-owned code, return vulnerability findings, and suggest patches for human review. Users do not receive unrestricted Mythos access through that interface.

Anthropic also works with security vendors to integrate Mythos into purpose-built products. The defender access model gives customers specific artifacts while keeping the underlying model behind controlled interfaces.

That mechanism has merit. It reduces the number of people who can directly test offensive prompts. It can also place model outputs inside established security workflows with permissions, logging, and human approval.

Yet a regulator cannot evaluate a restricted model solely by examining curated outputs. Independent testing requires freedom to challenge assumptions, explore edge cases, and reproduce harmful behavior. Those activities resemble the probing that Anthropic’s controls are designed to limit.

This creates a structural conflict. Anthropic wants to reduce exposure by controlling the model and its interfaces. ENISA needs meaningful access precisely because controlled demonstrations cannot establish the full risk.

The company’s model card and internal tests can provide useful evidence. They cannot substitute for evaluation by institutions with different incentives. A developer chooses its benchmarks, test environments, release thresholds, and interpretation of ambiguous results.

Public agencies bring a separate mandate. Their task is not to improve a product or protect a commercial release schedule. They assess systemic risks, compare developers, and connect technical behavior with public infrastructure obligations.

Mythos makes that independence especially important. The model is intended for cybersecurity work where a successful output can have immediate operational consequences. A finding might expose a widely used software component, while an exploit could threaten organizations before a patch exists.

The classified-system test illustrates both the promise and the uncertainty. An anonymous United States official said Mythos identified vulnerabilities in sensitive government systems within hours.

The official distinguished finding weaknesses from successfully exploiting them. That distinction matters because vulnerability discovery and reliable intrusion are different capability thresholds. Public discussion can blur them, especially when dramatic claims travel faster than technical evidence.

Senator Mark Warner described the test more strongly during a June hearing. He said the tool broke into almost all tested classified systems within hours, attributing that information to the head of the NSA and Cyber Command.

The NSA and Anthropic declined to comment on the reported exercise. Without detailed methodology, independent observers cannot determine the systems tested, the model’s autonomy, human assistance, or success criteria.

More than 100 cybersecurity experts and industry leaders later challenged the idea that Mythos was uniquely capable. They said other foundation and open-source models could also support security audits and training.

That disagreement strengthens the case for independent evaluation. If Mythos is uniquely dangerous, regulators need current access to understand its risks. If its abilities resemble competing systems, highly selective access may provide less safety than Anthropic suggests.

The Anthropic Mythos impact cannot be measured through reputation alone. It requires repeatable testing across comparable models, environments, safeguards, and versions. ENISA’s work can contribute to that evidence, provided its access remains technically meaningful.

Europe Has Access, but Not Equal Visibility

The agreement narrows the transatlantic access gap without establishing equal visibility into Anthropic’s newest systems.

Anthropic expanded Project Glasswing in June to approximately 150 additional organizations across more than 15 countries. That expansion followed an April preview involving a much smaller initial group.

Soon afterward, the United States government imposed temporary restrictions affecting foreign access to Mythos 5 and Fable 5. Anthropic disabled access broadly because it said it could not immediately enforce the directive at the required level.

The restriction applied even to foreign nationals working for Anthropic, according to the company’s account. United States authorities reportedly cited a potential security issue. Anthropic disputed whether the government’s response was warranted.

Access was later restored for approved United States organizations. However, the episode showed how quickly national policy can interrupt a private access program. It also gave European officials a concrete reason to discuss dependence on a technology controlled abroad.

The concern resembles a kill switch problem. European organizations can build security processes around an American model, then lose access through a government decision beyond their control. Even temporary disruption can matter during active incident response.

Granting Mythos 5 to ENISA reduces one part of that dependency. European evaluators can now generate their own evidence instead of relying entirely on Anthropic or United States agencies. They can test the model against European priorities and infrastructure assumptions.

The agreement does not eliminate version inequality. Anthropic’s public information says Mythos 5.1 remains limited to a set of United States organizations. The company is working to expand access but has not published a firm European schedule.

This pattern is visible elsewhere. Britain’s AI Security Institute reportedly received access to an earlier Mythos system but not Mythos 5.1 before release. The UK testing gap prompted officials to warn against fragmented international evaluation.

Anthropic had previously worked closely with the British institute. The reported exclusion therefore signals more than a routine scheduling delay. It suggests that access can narrow even for established government partners.

OpenAI provides a useful comparison without becoming the article’s main conflict. European officials say ENISA received GPT-5.6 Cyber and GPT-6 Astra. Britain also said its institute tested GPT-6 Astra before public release.

Those arrangements create pressure on Anthropic. Regulators can compare not only model capabilities, but also developer cooperation. A company’s evaluation policy can influence whether governments trust its broader safety commitments.

The comparison still requires caution. Public reporting does not reveal whether ENISA received identical access conditions across models. Testing duration, tool permissions, model snapshots, monitoring, and safeguard settings can all affect results.

Model names alone do not establish evaluation parity. ENISA needs enough technical detail to explain what it tested and how closely that system matches production deployments. Without such disclosure, readers cannot compare access across companies.

European authorities also face their own credibility test. Receiving a model does not guarantee a rigorous or timely evaluation. ENISA must show that it has appropriate infrastructure, specialist staff, and secure procedures for sensitive testing.

Its findings will need careful handling. Publishing exploit details too early could increase risk. Releasing only broad conclusions could prevent outside researchers from assessing the evidence.

A useful evaluation should distinguish several layers. It should separate vulnerability discovery from exploit creation, autonomous action from human assistance, and laboratory success from performance on maintained systems.

It should also document the safeguards present during testing. Mythos 5 differs from Fable 5 largely because selected cyber safeguards are lifted. Testing a constrained interface would reveal less than testing the access provided to operational partners.

For European companies, this is not a distant policy debate. Security teams may eventually receive Mythos-generated findings through integrated products. They will need to judge the reliability, provenance, and urgency of those findings.

Organizations should preserve the surrounding evidence in a searchable knowledge base. A model-generated alert gains value when reviewers can connect it with code changes, prior incidents, and human decisions.

The broader lesson is that model access and model sovereignty are not identical. ENISA can test an American system without controlling its availability, update cycle, or operating terms. Europe gains visibility, but Anthropic retains the gate.

Mythos 5.1 Keeps the Most Important Question Open

The largest uncertainty is whether Anthropic will give independent evaluators timely access to each new Mythos generation.

Anthropic describes Mythos 5.1 as its newest Mythos-class model, with gains in cybersecurity and biology. The company limits it to vetted organizations because both capability areas can support beneficial research or serious harm.

Anthropic’s Mythos 5.1 access page says current availability remains limited. It also says a Cyber Verification Program will include Mythos access in the future, without providing a fixed date.

That wording leaves three unresolved questions. First, it does not identify which international regulators will qualify. Second, it does not promise pre-release access. Third, it does not define what level of direct interaction accepted evaluators will receive.

These details determine whether the program produces meaningful oversight. Access after release supports research, but it cannot influence the original deployment decision. Limited interfaces can reveal defensive value while hiding offensive behavior.

There is also a version-matching problem. Anthropic says Fable 5.1 and Mythos 5.1 share an underlying model. Their safeguards differ, particularly for cybersecurity and biology tasks.

An evaluator cannot assume that results from Fable describe Mythos. Safety interventions can block tasks, redirect them to another model, or change the observed completion rate. The underlying weights are only one part of the deployed system.

Likewise, Mythos 5 results cannot settle questions about Mythos 5.1. A newer model may follow longer attack chains, use tools differently, or recover from failed attempts more effectively. It may also behave more safely under monitoring.

Anthropic publishes system cards and benchmark explanations, which improve visibility into its internal assessment. However, benchmark scores need context. Test design, scaffolding, trial counts, and safeguard interventions can materially change the result.

Cybersecurity benchmarks also risk becoming stale. Once tasks and vulnerabilities are widely known, training contamination becomes harder to rule out. Performance on a benchmark may not predict success against unfamiliar production systems.

Operational testing introduces another complication. Real networks contain identity controls, logging, segmentation, and defensive tools. A model that succeeds in an isolated range may struggle when those layers interact.

Conversely, a model can create risk without completing an autonomous intrusion. It might accelerate reconnaissance, suggest attack paths, or help a human operator iterate faster. Evaluations should measure these intermediate advantages.

Anthropic’s restricted distribution strategy therefore deserves neither automatic approval nor automatic rejection. The strategy addresses a genuine misuse problem. It also concentrates evidence and decision-making within the company and selected government relationships.

The skeptical test is straightforward. If controlled access is the central safety mechanism, Anthropic should show that independent evaluators receive current models under useful conditions. Otherwise, the mechanism protects the model from scrutiny alongside misuse.

ENISA’s Mythos 5 assessment can establish a baseline. Its value will increase if the agency publishes methodology, explains safeguard settings, and separates confirmed observations from vendor claims.

The evaluation will carry less weight if it appears after another major update without a clear bridge to the latest model. Timeliness should become part of the safety result, not an administrative footnote.

Three Signals Will Show Whether Access Becomes Oversight

The next test is whether this handover becomes a repeatable evaluation process instead of a one-time concession.

The first signal is access to Mythos 5.1. Anthropic does not need to provide unrestricted public availability, but ENISA needs a current evaluation target. Prompt access would strengthen the case that the three-month delay reflected transitional constraints.

Continued exclusion would weaken that explanation. It would show that European oversight remains tied to older releases while selected United States organizations receive newer capabilities.

The second signal is ENISA’s evaluation methodology. The agency should identify the tested version, access conditions, tool permissions, safeguards, and broad task categories. It should distinguish vulnerability discovery from successful exploitation.

Useful reporting does not require publication of dangerous technical details. It does require enough information to assess whether the testing examined genuine Mythos behavior. A vague statement that the model was evaluated would add little accountability.

The third signal is a durable international access framework. Anthropic’s Cyber Verification Program could provide that structure if it sets clear criteria, predictable review periods, and consistent treatment for qualified public evaluators.

A dependable framework would support Anthropic’s claim that restricted access can expand safely. Another sequence of case-by-case negotiations would reinforce concerns about selective and politically fragile oversight.

Developers and enterprise buyers should watch these signals because evaluation access shapes downstream trust. Security leaders may receive Mythos findings through products without ever interacting with the model itself.

They will need evidence about false positives, missed vulnerabilities, data handling, and human review. They should also ask whether access can disappear during a policy dispute or model transition.

The Anthropic Mythos EU access agreement is therefore a start, not a resolution. Europe finally has the June model while Anthropic controls the September successor.

Watch the next handover closely. If ENISA receives Mythos 5.1 promptly and publishes a credible testing framework, delayed access can become sustained oversight. If not, frontier AI evaluation will remain trapped behind the product cycle it is supposed to govern.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page