top of page

Anthropic Cyber Verification Program Expands Access, but Removes Key Claude Safeguards

7 hours ago
14 min read

Anthropic has expanded access to three advanced Claude models, despite their ability to complete complex offensive cybersecurity tasks when normal safeguards are removed. The Anthropic Cyber Verification Program now covers Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and future models. Approved organizations can use them for work that ordinary Claude users would see blocked.

The change opens a controlled route to AI capabilities that Anthropic has deliberately withheld from general access. Security teams can investigate malware, find vulnerabilities, and conduct authorized penetration tests. A smaller group can test systems tied to power grids, flight operations, telecommunications, banking, and government services.

This is more than a product-access announcement. Anthropic is testing whether identity checks, organizational controls, monitoring, and government review can replace restrictions built into the model experience. OpenAI has pursued its own phased access for advanced cyber models, while U.S. officials have taken a larger role in deciding who can use frontier systems.

The central question is no longer whether AI can assist cyber defenders. Anthropic says its earlier partners found more than 134,000 verified vulnerabilities with Claude models and related scanning work. The harder question is whether a provider can selectively remove safeguards without creating a privileged path that attackers eventually exploit.

Anthropic Cyber Verification Program Opens Three Access Levels

Anthropic is replacing one broad safety boundary with three increasingly permissive access levels.

The expanded verification program separates approved users into Defense Access, Red Team Access, and Specialized Access. Each level allows a different range of behavior and applies different eligibility requirements.

Defense Access supports work such as incident response, security operations, malware analysis, vulnerability validation, and defensive research. Eligible applicants can include companies, universities, nonprofits, government bodies, open-source maintainers, and individual researchers with credible disclosure records.

The program also includes critical infrastructure operators of different sizes. Anthropic specifically identifies organizations such as regional hospitals and municipal utilities. The company says it expects many legitimate defensive teams to qualify and aims to answer these applications within several days.

Red Team Access goes further. It permits authorized penetration testing and adversarial exercises against systems that an organization owns or has permission to test. In-house red teams, government teams, security consultancies, and penetration-testing firms can apply.

Organizations at this level must satisfy additional eligibility checks and security controls. Anthropic expects reviews to take several weeks. Applicants can receive Defense Access while the company considers their requests for the more permissive tier.

Red Team Access still has limits. Claude should block actions associated with physical harm or widespread disruption. Anthropic lists ransomware deployment, damage to physical systems, and testing high-risk safety systems among the restricted activities.

Specialized Access removes the largest number of cyber restrictions. It is reserved for organizations authorized to test systems where failure could endanger lives or disrupt major markets.

Those targets include flight operating systems, electric grids, telecommunications networks, interbank transfer infrastructure, and government administrative systems. Anthropic says it reviews every Specialized Access organization in depth with the U.S. government.

Existing Project Glasswing members will move into this level without seeking new approval for current models. Project Glasswing is Anthropic’s controlled initiative for giving selected infrastructure providers and security organizations access to its strongest cyber capabilities.

The program therefore relies on more than a verified email address or a standard enterprise contract. It links model access to the applicant’s identity, institutional role, testing authority, technical controls, and intended targets.

Organizations must also accept monitoring conditions. Anthropic generally requires data retention so it can detect possible misuse. Existing users of Fable 5.1 or Mythos 5.1 with zero-data-retention arrangements can temporarily keep those protections while joining the program.

Anthropic plans another option called Enterprise Frontier Safeguards. The company says this system will combine misuse controls with cloud storage managed by the participating organization. Until that system arrives, data governance remains a meaningful tradeoff for teams handling sensitive code or incident evidence.

The most important distinction is authorization. The same technical action can represent defensive validation or an unlawful intrusion, depending on the target and the operator’s permission. The program attempts to establish that context before loosening Claude’s restrictions.

That design creates the article’s core tension. Anthropic is not claiming that high-risk cyber behavior has become safe for everyone. It is claiming that verified institutions can receive stronger capabilities under a controlled access system.

Claude’s Cyber Capabilities Are Already Producing Results

The expansion follows evidence that advanced Claude models can compress months of vulnerability work into much shorter investigations.

Anthropic says Project Glasswing partners found at least 129,000 verified software vulnerabilities between April and July 2026. Its separate open-source scanning work identified another 5,500 verified vulnerabilities between April and October.

More than 33,000 of those findings received critical or high-severity ratings. The company says the total probably understates the program’s effect because its figures include reports from only 33 partners.

The data also has important limitations. Participants used different methods to confirm and categorize findings. Fewer than half disclosed how many vulnerabilities had been patched, often because remediation remained underway.

Finding a vulnerability is not equivalent to fixing it. Security teams must reproduce the issue, assess its severity, identify affected software, notify maintainers, create a patch, test the repair, and deploy it safely.

A model can accelerate discovery while making those later stages harder to manage. If automated systems produce findings faster than humans can verify them, security teams face a larger triage queue and a longer period of unresolved exposure.

Anthropic acknowledged that bottleneck during its earlier Glasswing expansion. The company said verification, disclosure, and patching had become more limiting than initial vulnerability discovery.

That pressure helps explain the broader access program. A small, centrally managed group cannot inspect every hospital network, open-source package, payment system, utility platform, or government application. Operators need direct access because they understand their own environments and can authorize realistic testing.

The potential gains extend beyond software scanning. Advanced models can reconstruct attack paths, analyze unfamiliar malware, generate potential patches, and repeat tests after defenders change a system.

A collaboration with Pacific Northwest National Laboratory illustrates the mechanism. Researchers created an agent scaffold, meaning a set of tools and instructions that allowed Claude to take controlled actions in a network environment.

The team used it to emulate attacks against a cyber-physical model of a water treatment facility. According to Anthropic’s infrastructure research, the system reconstructed an attack in three hours rather than several weeks.

The test used Claude Sonnet 4, an older model. During one run, a prepared method for bypassing Windows User Account Control failed. Claude detected the failure and selected another known technique to continue the simulated attack.

That adaptability is valuable to defenders because real environments rarely behave exactly like a written procedure. It is also the reason access becomes dangerous. A model that can recover from failed steps can help an attacker move beyond static instructions.

The new program aims to give legitimate teams enough freedom to reproduce that behavior. Basic code review and patching remain available through generally accessible Claude models. More interactive, multi-stage operations require verification because they resemble real intrusion activity.

This boundary will remain difficult to define. Vulnerability analysis often begins as benign code inspection and develops into an exploit chain. An incident responder may need to reproduce an attacker’s technique before blocking it.

Anthropic’s answer is to evaluate the operator and target, rather than judge every request only from its text. That shift puts identity, governance, and legal authorization alongside model-level filtering as core safety mechanisms.

Removing Cyber Safeguards Changes What Claude Can Complete

Anthropic’s own tests show that access tiers do not merely reduce inconvenience; they determine whether Claude can finish offensive operations.

The company evaluated Claude Opus 5.5 with CyScenarioBench, a benchmark for planning and executing multi-stage cyber operations under realistic constraints. It ran five attempts across each of ten challenges at every access level.

Without program access, Claude blocked every task at the first prompt. The generally available experience therefore prevented the evaluation from progressing, regardless of whether the operator had benign intentions.

Defense Access allowed slightly more activity, but safeguards still stopped 46 of the 50 trials before completion. Four tasks succeeded. That behavior fits a tier designed for investigation, response, and vulnerability work rather than full attack simulation.

Red Team Access produced a different result. No safeguard blocks occurred, and Claude completed 34 of 50 trials. Its 68 percent completion rate was effectively identical to the 67.6 percent result achieved without safeguards.

Those numbers clarify what Anthropic is granting. Approved red teams are not receiving the standard Claude experience with a larger request quota. They are receiving access to model behavior that closely approaches its unrestricted performance on this evaluation.

Specialized Access goes further by allowing authorized testing of safety-sensitive systems. It is meant for scenarios where even ordinary red-team access would retain restrictions because a failed test could affect physical operations, public services, or financial markets.

This structure is a calculated tradeoff. Strong universal filters can protect against misuse, but they can also stop defenders from rehearsing the attacks they must withstand. Broadly removing those filters would increase access for both groups.

Anthropic is betting that institutional verification can separate them. Defense users get broad eligibility with persistent restrictions. Red teams face deeper review but receive fewer blocks. A small set of organizations gains specialized access through joint scrutiny with the government.

Safety classifiers remain central to the system. A classifier is a separate model or control layer that evaluates requests and responses for prohibited behavior. It can interrupt an interaction even when the underlying Claude model is capable of continuing.

Anthropic has described how these controls divide cyber requests into categories, including clearly prohibited activity, high-risk dual-use work, lower-risk dual-use work, and ordinary defensive tasks. Its published classifier framework also recognizes that malicious and defensive operators sometimes use nearly identical techniques.

That ambiguity limits what automated filters can infer from a prompt. A request to establish persistence, extract credentials, or evade detection looks dangerous without reliable information about the target and authorization.

The verification program supplies that missing context before a session begins. It associates a user with an approved organization, an access level, contractual obligations, monitoring, and a defined range of permitted systems.

However, verification does not eliminate technical risk. Accounts can be compromised. Employees can exceed their authority. Contractors can lose access without permissions changing promptly. Authorized infrastructure can connect to unapproved systems.

Monitoring can detect suspicious use, but detection may occur after a model has generated a useful attack path. Rate limits and target controls can narrow the exposure, yet skilled operators often distribute work across tools and accounts.

The benchmark also evaluates only a sample of structured scenarios. A 68 percent completion rate does not establish how Claude will behave across unfamiliar software, custom industrial equipment, or chained attacks involving social engineering.

Anthropic’s evidence supports a narrower conclusion. Access controls can produce meaningfully different behavior across tiers, and advanced Claude models can complete many cyber tasks when restrictions are relaxed.

Whether those controls remain reliable at scale is not established. That question becomes more important as the applicant pool grows beyond the original Glasswing cohort.

Government Review Adds Security and Concentrates Authority

The U.S. government is becoming part of the access-control system for frontier cyber models, not merely an outside regulator.

Anthropic says it collaborates with the government when reviewing every Specialized Access organization. That arrangement gives public officials a role in deciding which institutions can use Claude against flight systems, power grids, banking infrastructure, and other sensitive targets.

The partnership has practical logic. Government agencies understand classified threats, national infrastructure, export controls, and the operational consequences of attacks on regulated systems. They can also confirm legal authority that a private model provider cannot assess alone.

Recent events show how far that relationship has developed. Anthropic reportedly worked with U.S. intelligence agencies to test a Mythos model against classified government systems.

An official told the Associated Press that the model identified some vulnerabilities within hours. The official stressed that identifying a weakness did not mean Mythos had exploited it in the same period.

Senator Mark Warner described the outcome more dramatically during a June hearing. He said the system had entered almost all the tested classified environments within hours, attributing that account to the leader of the National Security Agency and U.S. Cyber Command.

The NSA and Anthropic declined to confirm details for the classified testing. The gap between public descriptions is a reason to treat broad claims cautiously.

Government involvement has also produced conflict. In June, the Trump administration subjected frontier model releases to national security review and restricted access to some Anthropic systems.

Anthropic temporarily withdrew Fable 5 and Mythos 5 after a directive concerning foreign-national access. The government later approved Mythos for limited deployment to selected cyber defenders and infrastructure providers.

OpenAI faced a similar review for GPT-5.6 Sol. Both companies initially limited their newest systems to small groups, turning access to commercial AI models into an issue of national security policy.

Supporters can argue that unusually capable cyber systems require unusually careful release processes. Model providers lack complete visibility into intelligence threats, while governments cannot independently develop every relevant frontier model.

Critics see a less accountable arrangement. Representative Lori Trahan argued that political appointees were deciding company by company who received access without a clear law, process, or oversight structure.

That criticism applies to the expanded Anthropic Cyber Verification Program. Involving the government can improve vetting for sensitive targets, but it also concentrates authority in a process whose criteria are not fully public.

Organizations outside the United States face another concern. Anthropic previously expanded Glasswing to partners in more than 15 countries, and many served populations beyond their home markets.

Critical software is global. Open-source maintainers, cloud providers, chipmakers, telecommunications vendors, and financial networks routinely cross national borders. A system shaped heavily by one government’s security priorities can create unequal access.

The strongest cyber models could therefore become strategic resources distributed through geopolitical relationships. That would pressure organizations to satisfy both a private provider’s rules and a government’s national security requirements.

The program needs clear appeals, consistent standards, auditable revocation procedures, and transparent reporting about exclusions. Without those elements, verification risks becoming an opaque licensing system for an increasingly important defensive technology.

Anthropic has not presented the program as permanent public policy. It describes controlled access as a bridge while stronger safeguards are developed. Still, temporary systems often establish operational precedents that later become difficult to replace.

The Defensive Advantage Depends on Patching, Not Discovery

Claude can give defenders an advantage only if organizations repair vulnerabilities before attackers reproduce the same findings.

Anthropic’s stated goal is to shift the balance toward defense. That objective sounds intuitive because infrastructure operators can inspect their own systems, deploy patches, and coordinate incident response before publishing technical details.

Attackers have different advantages. They can choose the weakest target, reuse successful techniques, operate across jurisdictions, and move faster than institutional approval processes.

Advanced models can amplify both sides. A defender can scan millions of lines of code and prioritize dangerous flaws. An attacker can use similar reasoning to find an overlooked path into an exposed service.

The timing gap determines who benefits. A vulnerability discovered privately and patched quickly improves security. A flaw reported into a months-long remediation queue remains useful to attackers, even if defenders found it first.

Anthropic’s vulnerability totals therefore measure discovery, not a complete defensive outcome. More than 134,000 verified findings represent substantial research output. They do not reveal how many affected systems remain exposed.

The company says fewer than half of participating partners disclosed patch figures. That missing denominator prevents an independent assessment of whether discovery is outpacing remediation.

A high volume of automated findings can also create operational problems. Maintainers need enough evidence to reproduce a flaw, understand its impact, and distinguish an exploitable issue from a theoretical weakness.

Poorly prioritized reports can overwhelm volunteer open-source projects. Detailed exploit information can raise disclosure risk if it moves through too many organizations before a fix exists.

The expanded program places more responsibility on applicants. Security teams need vulnerability-management processes, isolated test environments, access logging, clear escalation paths, and authority to deploy repairs.

They also need durable records. AI-assisted investigations can generate long chains of hypotheses, tool outputs, code changes, and decisions. A searchable engineering knowledge base can help teams preserve that context without treating the model’s conclusions as verified facts.

Human review remains necessary. Models can misunderstand system boundaries, propose unsafe fixes, or identify patterns that fail under real operating conditions. Industrial systems add physical constraints that ordinary software benchmarks do not capture.

Specialized Access raises these stakes. A mistaken action against a flight-control environment, grid-management system, or interbank network can have consequences beyond data loss.

Anthropic says real-time restrictions continue at lower levels for activities that could cause physical harm or mass disruption. Specialized users receive fewer blocks precisely because their work sometimes requires testing those scenarios.

The program must therefore judge more than whether an organization appears legitimate. It must assess whether the organization can isolate targets, constrain model actions, supervise tests, recover from failures, and report anomalies.

Government review may support that assessment, but operational discipline remains local. A model provider cannot supervise every command inside a utility laboratory or a classified network.

Competition adds pressure. OpenAI has also made advanced cyber models available through controlled programs. If providers compete on which system completes more offensive tasks, access restrictions can become a commercial disadvantage.

If they compete on safety alone, defenders may choose models that offer fewer blocks and better attack simulation. That creates an incentive to relax controls while describing verification as the compensating measure.

Anthropic’s tiered structure is a credible attempt to manage that pressure. It does not resolve the underlying conflict. The same capabilities that make Claude useful against sophisticated attackers make any failure in verification more consequential.

The defensive advantage will be visible in remediation results, not model benchmarks. Patch rates, time to repair, recurrence rates, and prevented incidents matter more than the raw number of vulnerabilities found.

Three Signals Will Show Whether Controlled Access Works

The next test is whether Anthropic can expand participation without weakening verification or flooding defenders with unresolved findings.

The first signal is enrollment quality. Anthropic says it can review many Defense Access applications within days, while Red Team Access may take weeks. Faster approval is useful only if identity, authority, and security checks remain consistent.

The company should disclose aggregate application numbers, approval rates, review times, revocations, and major reasons for denial. It does not need to identify sensitive participants to show whether the system scales responsibly.

A sharp rise in access without corresponding monitoring capacity would weaken Anthropic’s case. Stable review standards and low misuse rates would strengthen it.

The second signal is the relationship between discovered and repaired vulnerabilities. Anthropic’s 129,000 partner findings and 5,500 open-source findings provide a baseline for discovery.

Future reporting should separate confirmed findings from duplicates, disputed reports, and vulnerabilities that affected obsolete software. It should also disclose how many issues received patches and how quickly those patches reached deployed systems.

Higher patch rates would support the claim that frontier models create a defensive advantage. A growing backlog of severe unresolved issues would suggest that automated discovery is exposing an organizational bottleneck rather than solving one.

The third signal is how Anthropic handles the first serious access failure. No verification system should be evaluated only during normal operation.

A compromised account, unauthorized target, classifier bypass, or accidental interaction with a live system would test the program’s response. The important measures would include detection time, containment, notification, revocation, and changes made afterward.

Transparent incident reporting would build confidence, even if some technical details remained confidential. Silence would make it difficult for outside organizations to assess the real risks of joining the program.

Enterprise Frontier Safeguards will also affect this signal. Anthropic says the planned system will combine privacy with misuse protection while letting eligible organizations keep information in their own cloud environments.

That arrangement could address concerns about sending sensitive code and incident data to a model provider. It could also make centralized monitoring harder if controls behave differently across customer-managed environments.

Competitor behavior belongs in the background, but it will shape Anthropic’s choices. OpenAI’s controlled cyber deployments will give buyers and regulators another model for comparing access, oversight, and incident response.

Broader access from a competitor would pressure Anthropic to accelerate approvals. A significant misuse event elsewhere could push it toward stricter requirements.

Readers should avoid treating the Anthropic Cyber Verification Program as proof that frontier cyber capabilities are now safe. It is a live governance experiment built around a difficult premise: known institutions can receive fewer restrictions because their identity and controls reduce risk.

That premise is testable. Anthropic has published benchmark results showing that its access levels change Claude’s behavior. It has also reported large numbers of verified findings from controlled deployments.

The missing evidence concerns operations at scale. We do not yet know how often legitimate requests are blocked, how many approved organizations violate rules, or how effectively Anthropic detects an account used outside its intended scope.

Developers and security leaders should watch the program because similar access systems can spread beyond cybersecurity. Frontier models with biological, financial, or autonomous research capabilities may also require permissions tied to verified users and approved environments.

Enterprise buyers should ask what an access tier changes technically. They should also examine data retention, monitoring, incident disclosure, employee authorization, and responsibility for model-generated actions.

For knowledge workers, the broader lesson is that access to advanced AI will not always arrive as a uniform product update. The most capable behavior may increasingly depend on who the user is, what institution supports them, and which systems they are authorized to test.

Anthropic has moved that future closer. Its program gives more defenders access to advanced Claude cyber capabilities while preserving stronger restrictions for everyone else.

The outcome will depend less on headline vulnerability counts than on disciplined remediation and accountable oversight. Will Anthropic publish enough evidence to show that verified access creates safer infrastructure, or will the first major failure expose verification as the weakest safeguard?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page