Anthropic Google and OpenAI Guardrails Are Slowing Offensive Cybersecurity Research
- Martin Chen

- 1 day ago
- 14 min read
Anthropic Google and OpenAI safety controls now face a difficult test: researchers say stricter guardrails are obstructing authorized offensive cybersecurity work. The restrictions target malicious hacking, yet the same controls can block vulnerability validation, reverse engineering, and exploit development.
That conflict became harder to ignore after several offensive security researchers described repeated refusals and inconsistent results. Their work involves finding unknown flaws before criminals exploit them, often by developing controlled proofs of concept.
The problem is not simply that AI companies decline dangerous requests. Researchers say the systems struggle to distinguish an authorized test from an actual attack. That ambiguity creates delays, limits reproducibility, and pushes sensitive work toward locally hosted open models.
OpenAI and Anthropic have introduced verification programs intended to reduce that friction. Google also restricts generative AI uses involving malware, infrastructure disruption, and bypassing safety filters. Together, these policies show how leading American AI companies are drawing boundaries around dual-use cyber capabilities.
The central dispute is sharper than a routine complaint about model refusals. Offensive and defensive security often require the same technical steps. A model cannot always identify the operator's authorization from the code, command, or vulnerability placed in its context.
That leaves frontier labs choosing between two costly errors. A permissive model can help an attacker move faster. An overcautious model can prevent a defender from understanding and fixing the same weakness.
Researchers Say Legitimate Cyber Work Is Getting Blocked
The immediate change is that cyber guardrails now interrupt practical research workflows, not only obviously malicious requests.
A July 23 cybersecurity investigation documented complaints from researchers who find vulnerabilities and develop exploits. Several said frontier models refused legitimate requests or produced inconsistent answers during authorized work.
Offensive security means testing systems from an attacker's perspective with permission from the owner. It includes penetration testing, exploit validation, red teaming, and some forms of reverse engineering.
These activities can expose weaknesses that a conventional code review misses. A suspicious code path becomes actionable only after a researcher establishes whether it can be reached and exploited.
Chris Anley, chief scientist at NCC Group, told TechCrunch that asking a model to exploit a bug can confirm whether it deserves remediation. A refusal at that stage does not merely remove convenience. It can interrupt the evidence-building process that helps a company prioritize a fix.
Anley compared the technology to a hammer, which can serve as both a tool and a weapon. His broader point concerns technical overlap. Instructions useful for exploitation may also be essential for demonstrating that vulnerable code creates real risk.
That overlap becomes especially important with zero-days, meaning vulnerabilities unknown to the affected vendor when discovered. Researchers frequently need to trace execution, manipulate memory, or construct unusual inputs before they understand a flaw.
Those steps resemble attacker behavior because researchers are reproducing the attack. A classifier sees suspicious terms, code, and commands. It may not see the contract or laboratory authorization that makes the work legitimate.
One researcher at a smartphone-component manufacturer told TechCrunch that Anthropic's tools became barely useful for vulnerability discovery outside its verification program. The researcher said the system stopped once it detected security-related work.
Chris Thompson, CEO of RemoteThreat, described another failure mode. He said guardrails behaved differently across sessions, including within vetted programs offering looser restrictions.
That inconsistency matters because vulnerability research depends on repeatable experiments. A researcher must determine whether a result comes from the target, the testing method, or the model's changing intervention.
When policy enforcement changes without a clear explanation, the model adds another uncontrolled variable. Time moves from analyzing the vulnerability toward rephrasing prompts and diagnosing refusals.
The restrictions do not affect every practitioner equally. Giuseppe Cali, who finds zero-days and develops exploits, told TechCrunch that guardrails had not impeded his work.
Cali uses AI for initial reverse engineering and for building supporting tools. He keeps actual vulnerability discovery and weaponization under his own control, partly because he enjoys that work.
His experience establishes an important limit on the criticism. Frontier models can remain useful when researchers assign them lower-risk supporting tasks. The conflict intensifies when a workflow approaches exploitability, payload behavior, or operational testing.
The current debate therefore concerns access to particular capabilities, not whether AI has any value in security. Models can summarize code, explain functions, propose tests, or assist with documentation without crossing the hardest policy boundaries.
Problems emerge when a researcher asks the model to connect those steps. That connection is often where a possible flaw becomes a verified security finding.
Why Anthropic Google Policies Cannot Cleanly Separate Attack From Defense
Anthropic Google and other frontier-lab policies confront a classification problem that prompts alone cannot reliably solve.
Anthropic says its real-time safeguards block prohibited use and high-risk dual-use activity. Its prohibited category includes conduct with little legitimate defensive value, such as ransomware development or mass data exfiltration.
The high-risk dual-use category is more complicated. Anthropic explicitly includes vulnerability exploitation and offensive security tool development, both of which can serve legitimate defensive purposes.
Those requests are blocked by default on affected Claude models. Verified users can apply for adjustments through Anthropic's Cyber Verification Program, or CVP.
Anthropic says approved practitioners may still encounter blocks. The company acknowledges that it can incorrectly decline eligible applicants and that approved users can face restrictions on legitimate work.
That acknowledgment reflects the core technical challenge. A model may receive identical exploit-development instructions from a penetration tester and a criminal. The visible difference is often authorization, not the requested operation.
Google's prohibited-use policy similarly bars generative AI content that facilitates malware, infrastructure abuse, or safety-filter circumvention. It allows exceptions where educational, scientific, or public benefits outweigh the harm.
The exception language recognizes context, but enforcing context at scale remains difficult. A user can claim to own a target without proving it. A legitimate researcher can also work on a product without controlling its infrastructure.
OpenAI describes cyber capabilities as inherently dual-use. It says defensive and offensive workflows rely on much of the same knowledge and many of the same techniques.
The company reported a steep rise in cyber performance during 2025. Its models' capture-the-flag results increased from 27% in August to 76% in November, according to its cyber resilience plan.
Capture-the-flag challenges are controlled security exercises with intentionally vulnerable targets. They test skills such as reconnaissance, reverse engineering, exploitation, and privilege escalation.
OpenAI uses higher capability thresholds to anticipate models that can develop remote zero-day exploits or assist complex intrusions. That trajectory explains why the company will not treat every cyber request like ordinary coding.
Google DeepMind has taken a similar risk-based approach at the model-development level. Its safety framework includes cybersecurity among the domains requiring capability evaluations and escalating mitigations.
These policies respond to a real concern. A capable model can compress specialized knowledge, automate repetitive work, and coordinate tools. The same efficiency available to a defender can lower the effort required for abuse.
However, aggressive prompt filtering is a blunt response to an identity and authorization problem. Technical content alone provides weak evidence about the user's purpose.
Simple wording can also be misleading. A benign prompt may contain terms such as shellcode, exploit, persistence, or credential extraction because those concepts appear in legitimate assessments.
A malicious request can use sanitized language. An attacker might describe credential theft as account recovery or disguise exploitation as compatibility testing.
Classifiers must evaluate code, intent, target, account history, and surrounding activity. Even then, they make probabilistic judgments rather than verifying legal authorization.
False positives become especially likely during exploratory research. Researchers do not always know what a suspicious code path will become. They may need to test several offensive hypotheses before identifying the actual vulnerability.
This uncertainty conflicts with systems that expect a neatly defined defensive purpose at the start. The research process often produces that explanation only after dangerous-looking technical work has begun.
Verified Access Helps, but It Does Not Remove the Friction
Verification programs improve access for known defenders, yet they cannot guarantee unrestricted or predictable research workflows.
Anthropic's CVP is an application-based program for legitimate high-risk cyber work. The company says it aims to issue a review decision within two business days.
Approval is connected to a specific organization. Researchers can still receive blocks if they use another workspace or encounter an activity that remains prohibited.
Availability also varies by platform. Anthropic states that CVP is not currently available through Amazon Bedrock or Google Vertex AI. Third-party applications using Claude may not participate either.
Those distinctions create operational complexity. A security team can receive different treatment depending on its cloud provider, workspace configuration, or software integration.
The program also requires data retention. Anthropic advises organizations using zero data retention to create a separate workspace where retention is enabled.
That condition raises a second issue beyond guardrails. Offensive research frequently involves confidential source code, unpatched vulnerabilities, proprietary firmware, and information that could enable compromise.
Sending those materials to a hosted model can create unacceptable exposure, even when contractual protections exist. Some researchers therefore avoid frontier services before a vendor has patched the flaw.
Paolo Stagno, chief technology officer at Crowdfense, told TechCrunch that his team uses frontier models for reverse engineering. It avoids using them to find vulnerabilities or create exploits.
Stagno cited concerns about leaking sensitive vulnerability information or having data absorbed into future training. For exploit work, his team uses locally operated open models that do not require sending the material to a provider.
OpenAI's trusted cyber access system uses several capability levels. Standard access supports common defensive tasks, while verified access reduces some restrictions for authorized work.
More specialized access covers offensive testing, including exploit development, penetration testing, reverse engineering, and red teaming. Approval for one access level does not automatically grant every cyber-specialized model.
OpenAI also says the program does not remove every safeguard or refusal. Access remains limited to approved users, internal workflows, and systems the organization owns or has permission to test.
Those conditions are defensible. A provider cannot safely interpret verification as blanket permission to attack any target.
Yet they also reveal why verification does not fully solve the research problem. Authorization is granular. It can apply to one system, one period, one test method, or one customer engagement.
A general account-level approval cannot capture every scope change. Continuous document review would introduce more delay and expose additional customer information.
OpenAI previously said its trusted program had reached thousands of verified defenders and hundreds of teams protecting critical software. That scale shows demand, but it does not measure false positives or abandoned workflows.
The missing performance indicators are practical ones. Researchers need to know how frequently approved requests get blocked, how long appeals take, and whether decisions remain stable between model updates.
They also need useful explanations. A generic safety message does not reveal whether the trigger was the target, code behavior, requested output, or accumulated account activity.
Without that information, users experiment with prompt wording. This behavior can resemble attempted circumvention even when the researcher only wants a legitimate result.
Better enforcement would focus on the entire operating environment. Providers can combine identity verification with isolated execution, target allowlists, rate limits, audit logs, and controlled network access.
That model resembles a cyber range, which is an isolated environment for security exercises. It offers stronger evidence of authorization than prompt classification alone.
It also places more responsibility on infrastructure. Providers must confirm that a supposedly isolated target cannot become a bridge to an external system.
This approach will not cover every researcher. Independent specialists and small consultancies may lack the organizational documentation expected by enterprise-focused programs.
Security research has long benefited from outsiders who investigate products without a formal relationship with the vendor. Restricting advanced access to large, easily verified institutions could narrow that community.
The result would be a two-tier system. Major companies receive specialized models and support, while independent researchers rely on consumer tools, local open models, or manual methods.
Guardrails Are Pushing Sensitive Work Toward Open Models
When hosted frontier models become unpredictable or unsuitable for confidential work, researchers gain a reason to run open models locally.
Several researchers told TechCrunch that they fall back on downloadable models when American frontier services refuse offensive tasks. Thompson specifically pointed to Chinese open models such as GLM.
Local deployment changes the control structure. Researchers can choose the model version, preserve prompts, disable external connectivity, and keep vulnerable code on their own hardware.
They can also reproduce an experiment after a vendor updates its hosted service. A fixed model checkpoint behaves more consistently than a service whose safeguards can change without notice.
This does not mean every open model matches the reasoning quality of a leading hosted model. Researchers must compare capability, context handling, hardware needs, and tool integration.
Open models also transfer more security responsibility to the operator. A poorly isolated agent can execute unsafe commands, expose secrets, or reach production systems.
Still, local control addresses two complaints at once. It removes provider-level refusals and reduces the need to send unpublished vulnerability data to an external service.
That combination matters more than benchmark leadership. A slightly weaker model can be more useful if it remains available throughout a sensitive, repeatable workflow.
The shift also creates a strategic tension for American AI companies. Strict safeguards can reduce misuse on their own platforms while redirecting legitimate experts toward systems outside their governance.
Those researchers then provide feedback, integrations, and workflow knowledge to another model ecosystem. Their usage can improve tools around models that have fewer restrictions.
This is not proof that guardrails make the internet less safe. Attackers can use open models regardless of what verified defenders choose.
However, defender migration can weaken the argument that hosted restrictions preserve an advantage for responsible users. The advantage matters only if those users can complete meaningful work.
OpenAI frames its strategy around giving defenders better tools while limiting malicious uplift. Anthropic similarly offers adjusted safeguards for verified professionals.
Both goals depend on calibrated access. A control that blocks nearly all dangerous-looking behavior can reduce abuse, but it can also eliminate the capability defenders were promised.
The migration pressure extends beyond explicit refusals. Hosted models can become unattractive when retention requirements, platform exclusions, or unclear appeals complicate confidential engagements.
A consulting firm may serve multiple customers with separate authorization boundaries. It cannot casually mix their source code and findings inside one retained workspace.
An independent researcher may investigate a widely deployed product before contacting its maker. That person cannot always supply a client contract proving authorization.
Bug-bounty programs create another gray area. They invite testing under published rules, but a model provider may not be able to verify whether each requested action stays within scope.
These situations expose the limits of company-level vetting. Trust attaches to a person or organization, while authorization attaches to a specific operation.
Locally operated models avoid that verification gap. They also remove the provider that could detect large-scale abuse, suspend access, or investigate suspicious patterns.
The tradeoff therefore moves rather than disappears. Hosted services provide oversight but can create friction and confidentiality concerns. Local models provide control but reduce centralized enforcement.
American frontier labs cannot reverse open-model availability through stricter refusals. Their more realistic option is to make responsible hosted access better than the local alternative.
That means predictable policies, faster appeals, meaningful privacy controls, and environments designed for controlled exploitation. Raw model capability alone will not keep security researchers on a platform.
The Researchers' Criticism Has Important Limits
Offensive researchers identify real workflow failures, but their interests do not settle how broadly dangerous AI capabilities should be released.
Some offensive security businesses discover, acquire, or sell vulnerabilities to government customers. Their work does not always lead to immediate disclosure and patching.
Mark Dowd, a prominent researcher quoted by TechCrunch, has sold zero-days to Western governments. He acknowledged that this background can shape his view of corporate restrictions.
Governments value undisclosed vulnerabilities because intelligence agencies can use them while targets remain exposed. That market complicates any simple equation between offensive research and public defense.
A model provider must consider more than whether the customer appears reputable. It must also consider whether assistance could expand surveillance, intrusion, or exploit stockpiling.
Legal authorization is not identical to public benefit. A government-approved operation can remain controversial or create systemic risk if a vulnerability affects widely used software.
Researchers also differ over how central AI should become. Cali's workflow suggests that useful assistance can stop before automated bug discovery or weaponization.
That approach preserves human judgment at the riskiest stages. It also reduces the chance that a model will transform an incomplete idea into a reusable attack method.
Meanwhile, the providers' risk argument is not hypothetical. OpenAI says model cyber performance has risen rapidly, and all three labs treat advanced cyber capabilities as a serious safety domain.
As models gain longer-running agency, a single response becomes less important than an action sequence. An agent can inspect code, generate tests, execute commands, evaluate failures, and revise its plan.
That capacity changes the stakes. A refusal that once blocked a short malware request may need to govern thousands of coordinated actions.
It also makes isolated testing essential. An agent operating with tools can cross boundaries that a text-only assistant cannot reach.
Researchers deserve predictable access, but providers need evidence that the surrounding environment will contain failures. Verification alone cannot supply that evidence.
There is also no public, standardized measurement of guardrail quality for legitimate cyber work. Anecdotes reveal failure modes, but they do not establish overall false-positive rates.
The TechCrunch interviews cover varied organizations and workflows. They provide credible warning signs, not a representative survey of the security industry.
Providers likewise publish limited data about approved users' experiences. Enrollment counts reveal reach, while saying little about task completion or model usefulness.
This evidence gap encourages both sides to overstate their case. Researchers can interpret a refusal as proof that safety policy is arbitrary. Providers can treat program availability as proof that legitimate access works.
A better evaluation would test realistic authorized workflows across multiple models. It should measure task success, inappropriate refusals, unsafe compliance, consistency, appeal time, and data-handling requirements.
The benchmark should include ambiguous cases. Easy defensive prompts and blatant ransomware requests do not test the contested boundary.
Scenarios could include exploit validation inside a cyber range, reverse engineering of malware, analysis of a bug-bounty target, and development of a safe proof of concept.
Independent evaluators would also need access to the highest-risk model tiers. Otherwise, they would measure public restrictions without examining whether verification actually resolves the problem.
The results should not disclose operational details that enable abuse. Aggregate reporting can still show whether controls distinguish legitimate work more accurately over time.
Until such evidence exists, strong conclusions remain premature. Guardrails clearly cause friction for some researchers, but removing them would create a different and potentially larger risk.
The practical objective is not unrestricted access. It is accountable access that remains useful under realistic offensive testing conditions.
What Anthropic Google and OpenAI Must Prove Next
The next phase should be judged by workflow reliability, controlled execution, and whether responsible researchers remain on governed platforms.
The first signal is measurable performance inside verification programs. Anthropic and OpenAI should report approval timelines, appeal outcomes, and false-positive rates for legitimate dual-use requests.
Those numbers need context, including model, access level, and task category. A single program-wide percentage could hide severe problems in exploit validation.
Anthropic already says it aims to decide CVP applications within two business days. The more important question is what happens after approval.
If verified researchers still encounter frequent unexplained blocks, vetting has shifted the boundary without solving the workflow problem. Falling false-positive rates would strengthen the case for calibrated safeguards.
The second signal is the spread of controlled research environments. Frontier labs can provide isolated workspaces with audited tools, constrained networking, and clear target authorization.
Such environments would let models perform dangerous-looking tasks without granting open access to external infrastructure. They would also make incident review more concrete than prompt-level speculation.
Success would require support across cloud platforms and third-party tools. Anthropic's current CVP availability gaps show how access can break when the model reaches users through intermediaries.
Privacy controls will matter as much as execution controls. Security teams need credible options for handling proprietary code and undisclosed vulnerabilities without creating additional exposure.
If companies pair cyber ranges with strong retention choices, more researchers can justify using hosted models. If retention remains mandatory, sensitive work will continue moving locally.
The third signal is researcher behavior. Providers should watch whether respected offensive teams use specialized frontier models for exploit validation, not just code summarization.
Continued migration toward GLM and other downloadable models would weaken claims that verified access gives defenders a practical advantage. Stable adoption would suggest safeguards are becoming usable.
Google belongs in this comparison even though the recent researcher accounts focused mainly on Anthropic and OpenAI. Its policies and frontier framework reflect the same underlying tradeoff.
The anthropic google keyword pairing also captures a broader market reality. Security teams compare governance, deployment, privacy, and access across providers, not only raw benchmark scores.
No company can solve dual-use classification with a better refusal message. The decisive improvements will combine identity, environment, authorization, monitoring, and a transparent path for correcting mistakes.
Developers and enterprise buyers should ask direct questions before adopting an AI model for security work. Which tasks trigger enhanced safeguards? Can approved users appeal during an active engagement?
They should also examine retention, regional processing, platform availability, auditability, and model-version stability. These details determine whether a product can support an actual security program.
Offensive researchers should document false positives without publishing material that enables harm. Comparable evidence will make it harder for providers to dismiss failures as isolated misuse.
AI companies, meanwhile, should treat legitimate refusals as safety defects. A system that blocks defenders indiscriminately does not achieve the intended balance.
The coming test is straightforward. Can Anthropic, Google, and OpenAI preserve meaningful oversight while giving authorized researchers dependable access to the capabilities attackers already seek?
If verified programs become consistent and confidential, governed frontier models can retain responsible experts. If not, those experts will keep moving toward local alternatives with fewer restrictions and less provider oversight.


