On Anthropic’s AI Misuse Report: Agents Shift the Cost to Defenders
Anthropic published its third AI misuse report after detecting Claude in cyberattacks, surveillance systems, influence campaigns, weapons research, and industrial model extraction. The evidence behind On Anthropic’s AI Misuse Report covers activity observed between December 2025 and August 2026. Its central warning is not that AI invented new malicious goals. It is that agents made existing operations cheaper, faster, and easier to scale.
The report describes people choosing targets, defining objectives, and reviewing consequential outputs. AI systems then handled reconnaissance, software development, vulnerability research, data processing, propaganda production, and parts of exploitation. Several operators also connected Claude to external tools and persistent memory, letting automated workflows continue across sessions.
That division of labor matters more than any isolated malicious prompt. It places AI inside the operating machinery of cybercrime and state surveillance. Attackers can now preserve expertise in software, repeat successful procedures, and run several workstreams at once.
Bruce Schneier highlighted the same pattern in his misuse report analysis. The meaningful shift is toward industrialized execution, with humans retaining authority while machines carry more of the operational load.
On Anthropic’s AI Misuse Report, the Workflow Is the Story
The report’s most consequential evidence concerns complete workflows, not isolated answers generated by a chatbot.
Anthropic’s threat intelligence report documents activity by criminal groups, commercial surveillance vendors, politically motivated individuals, and state-aligned organizations. The cases span cyber operations, propaganda, surveillance, conventional weapons, biological research, and unauthorized model distillation.
The company says it selected notable or novel incidents, not a representative sample of everything happening on Claude. That distinction limits conclusions about prevalence. However, the cases still reveal what motivated operators can assemble when models connect to code, browsers, scanners, databases, and other tools.
Some operators ran agent swarms, meaning a coordinating model divided a large objective among several specialized AI agents. Those agents worked in parallel on reconnaissance, exploitation research, and post-compromise tasks. Persistent campaign memory stored target lists, credentials, instructions, and engagement status between sessions.
In one reported operation, automated workflows analyzed appliance firmware and binaries with decompilation tools. The system formed vulnerability hypotheses, wrote proposed exploits, and tested them against laboratory copies of targeted products. Anthropic says one continuous workflow produced more than a dozen possible zero-day findings within one month.
A zero-day is a previously unknown software vulnerability without an available defensive patch. The report does not establish that every suspected finding was valid or used against real systems. It does show an automated process iterating through work that traditionally required sustained specialist attention.
Other agents repeatedly searched exposed internet infrastructure. They fingerprinted services, compared entry points with known vulnerabilities, and added new findings to persistent project memory. A separate fleet of 13 standing agents collected public material from military, government, policy, and social sources.
This architecture changes the economics of malicious work. A human no longer needs to perform every search, test, translation, or documentation task personally. The operator can supervise an expandable collection of automated workers and concentrate on target selection.
That is why a list of individual prompts understates the development. Each prompt might appear ordinary when reviewed alone. The danger emerges from orchestration, accumulated context, tool access, repetition, and the ability to act on outputs.
Anthropic says it banned linked accounts and used its findings to improve safeguards. Yet account removal addresses access to one provider, not the underlying expertise, code, stolen credentials, or infrastructure. Operators can preserve those assets and move parts of a workflow elsewhere.
The report therefore turns AI misuse from a content-moderation problem into an operational-security problem. Blocking harmful text remains useful, but defenders must also detect automated behavior across identities, sessions, tools, and compromised accounts.
AI Agents Are Compressing the Attack Cycle
AI agents do not remove humans from cyberattacks, but they let fewer humans operate across more targets and more technical layers.
One campaign described by Anthropic used AI across reconnaissance, initial access, collection, exfiltration, and persistence. The actor reportedly fingerprinted email and remote-access systems, built target lists, and created infrastructure for device-code phishing.
Device-code phishing abuses a legitimate login process to persuade a victim to authorize an attacker-controlled session. Once the attacker obtains that access, the session can expose cloud email and other connected services. The technique often looks less suspicious because it uses a real authentication flow.
Anthropic says the operation accessed mail records from at least eight organizations. The targets included a national prosecutor’s office, a military education institute, and a regional intergovernmental organization. Another intrusion exposed more than 300,000 national identity records and registry information for over 500,000 companies.
The AI system helped organize hundreds of gigabytes of stolen information, according to the report. It also assisted with credential harvesting, lateral movement, and adapting malicious software after security products detected earlier versions.
This closes part of the feedback loop between attacker action and defensive response. A detection no longer creates the same delay if software can analyze the failure, revise an artifact, and prepare another version quickly. Human approval can remain in the loop while machine iteration accelerates.
A different actor used an autonomous exploitation pipeline against production web applications. Worker agents tested injection flaws, authentication bypasses, cross-site scripting, and server-side request forgery without continuous human supervision. Potential credentials and findings flowed back into a shared workspace.
Server-side request forgery is a flaw that tricks a server into making requests on an attacker’s behalf. It can expose internal systems that are otherwise unreachable from the public internet. Automating such tests raises the number of endpoints an operator can probe.
The report also describes fraud infrastructure that automated account registration, inbox monitoring, CAPTCHA solving, and identity-verification steps. Another workflow placed a reverse proxy between victims and genuine identity checks, capturing authenticated sessions and submitted documents.
These examples show why On Anthropic’s AI Misuse Report matters to ordinary organizations. A company does not need to be the attacker’s original target to suffer damage. Its exposed keys, cloud tenants, SaaS applications, suppliers, or customer records can become intermediate resources.
One French-speaking attacker reportedly targeted 42 political and affiliated organizations. The actor gained internal access to at least 14 and removed an estimated 12 to 26 gigabytes of database material. A compromised campaign platform exposed about 140,000 records, including political opinions.
Independent reporting about the French campaign connected several details with affected organizations. That reporting offers valuable corroboration, although Anthropic remains the primary source for much of the technical account.
The actor also developed a previously undocumented WordPress exploitation technique and succeeded against at least four websites, Anthropic says. Other tooling harvested submitted credentials, poisoned backups, and placed remote-access code inside website assets.
This was not an entirely autonomous attack. The human chose targets, operated infrastructure, interpreted results, and pursued political objectives. The important change was leverage: one person could build and maintain a campaign with the breadth of a small technical team.
The Main Conflict Is Attacker Scale Versus Defender Accountability
Attackers can distribute work across agents, but defenders still carry the full legal, operational, and financial consequences of every failure.
Traditional security planning often assumes that offensive work consumes scarce expertise. Skilled attackers must study infrastructure, create tools, inspect stolen data, and adjust their methods after detection. Those constraints limit how many targets a group can pursue at once.
Agentic systems weaken those constraints. They can run reconnaissance continuously, document results consistently, translate material, and retry technical tasks. They also preserve procedural knowledge that would otherwise live in one operator’s memory.
Defenders face a less forgiving equation. Every exposed credential, forgotten application, vulnerable vendor, or excessive permission can become an entry point. Attackers need one viable path, while defenders must protect a changing collection of systems and identities.
The AI supply chain itself has also become a target. Anthropic describes fake services that promised discounted access to frontier models but installed credential harvesters. Those programs stole account credentials and authenticated session tokens from customers’ devices.
The attackers then sold or reused that access through fraudulent proxy networks. Some services silently routed customers to a different model while continuing to collect new credentials. This gave criminals a renewable source of legitimate-looking accounts after earlier keys were revoked.
Another operation tested 30 targets while seeking access to a prerelease Claude model. Anthropic says every attempt failed and its own systems were not compromised. The stolen keys involved belonged to customers whose environments had already been breached.
That distinction is important. A provider can secure its central infrastructure while attackers exploit weaker points around it. Developer laptops, browser sessions, automation platforms, third-party routers, and copied configuration files all become part of the effective security boundary.
Enterprises should therefore treat AI credentials like other privileged production secrets. Keys need limited scopes, short lifetimes where possible, usage monitoring, and reliable revocation. Unusual model traffic may reveal a compromised environment before an attacker reaches a more obvious objective.
Security teams also need visibility into agent behavior. A single account launching persistent scans, parallel tool calls, or repetitive extraction jobs presents a different risk from an ordinary conversation. Detection should consider sequences of actions rather than one prompt at a time.
This creates pressure on Anthropic, OpenAI, Google, Microsoft, cloud providers, and model-routing services. Each sees a different portion of the chain. No participant can reconstruct the whole campaign without exchanging useful threat information.
However, more monitoring introduces its own risks. Providers may inspect prompts, outputs, files, tool calls, and account relationships to identify abuse. Those practices can expose sensitive customer information or create broad surveillance capabilities inside the service itself.
The resulting tradeoff is not simply safety versus privacy. Weak monitoring can let coordinated attacks continue. Excessive monitoring can erode confidentiality, produce false accusations, or place valuable customer data into another centralized system.
Providers need bounded retention, restricted investigator access, clear escalation rules, and meaningful transparency about automated enforcement. Customers also need to know what telemetry exists and when information may be shared with outside organizations.
On Anthropic’s AI Misuse Report offers unusually detailed visibility into detected activity. It does not provide a complete measure of false positives, undetected campaigns, or the privacy cost of the investigations. Those unanswered questions belong beside the operational findings.
Surveillance and Propaganda Show the Same Labor Substitution
The report’s surveillance cases show AI replacing parts of an engineering and analytical workforce without changing whom powerful institutions choose to target.
A consultant working for national security authorities in Mali reportedly used Claude to engineer a mass-interception platform. The system could collect communications data from the country’s mobile operators and generate dossiers on selected people.
Anthropic says the model helped design the underlying software rather than analyze the final dossiers. According to separate surveillance reporting, the platform was eventually deployed locally with other models after Anthropic intervened.
That outcome exposes a limit of provider enforcement. A cloud service can close accounts and make one workflow harder to run. It cannot reliably erase exported code, technical knowledge, collected data, or access to alternative models.
In Iran, actors used Claude to build a malicious Firefox extension that harvested identities from social networks. Another Iranian unit analyzed hundreds of thousands of posts and identified 39 opposition accounts for monitoring, according to Anthropic.
The report describes a Chinese religious-affairs intelligence operation that had reportedly contracted from many analytical teams to one office. With AI assistance, that office produced thousands of investigations per month.
Other actors used Claude to score articles and social posts for political sensitivity. They assembled locations, demographic attributes, political views, and confidence ratings into structured records. These outputs could support questioning, monitoring, or other forms of control.
One PRC-aligned operator lacking Arabic skills allegedly ran a multiday recruitment operation targeting Uyghurs in Syria. Claude drafted outreach in a regional dialect, translated replies, role-played quality checks, and formatted results for a suspected case officer.
The model did not choose the persecuted community. Human institutions supplied the target, mission, and intended use. AI reduced the language, staffing, and technical barriers that might otherwise slow the operation.
That pattern also appears in influence campaigns. Anthropic details nine cases originating across Russia, Iran, Turkey, the Gulf, South Asia, Africa, and Europe. The campaigns targeted audiences on six continents.
Operators built fake profiles, news sites, political messages, and coordinated posting plans. Some campaigns aligned with elections. Russian state media produced fabricated claims before Moldova’s September 2025 vote, while a Kenyan operator prepared fake grassroots content before the 2027 election.
AI did not invent deceptive political communication. It helped produce personas, translations, articles, targeting strategies, and variations at lower marginal cost. A small team could maintain a larger facade of independent public participation.
Anthropic has an unusual observation point because it can see campaigns while operators are planning them. Social networks often encounter the operation only after content appears publicly. Model providers may detect target selection and content creation earlier.
Their visibility then declines when content leaves the model platform. Anthropic says it uses open-source research, partner data, and public reporting to understand downstream activity. That means some planned campaigns may never launch, while others may migrate beyond its view.
Readers should resist treating every detected prompt as proof of a completed real-world operation. Intent, preparation, deployment, reach, and measurable impact are different stages. The report is strongest when it connects model activity with infrastructure or corroborating evidence.
Still, the direction is clear. AI is entering the routine bureaucracy of surveillance and political influence. It can help institutions process more people, maintain more personas, and prepare more reports without proportionally expanding staff.
Weapons Research Raises a Harder Safeguard Problem
Anthropic’s findings move the debate beyond malicious writing assistance and toward software that connects analysis with physical systems.
The company says it disrupted six conventional-weapons cases involving three actors in China, two in Russia, and one in Yemen. Four involved software for weapons hardware, while two focused on procurement or intelligence collection.
The reported projects included guided rockets, missile guidance, drone-swarm software, torpedo interception, electronic-warfare targeting, and air-defense suppression. One guided-rocket program included a live field test, according to Anthropic.
The actors generally had relevant hardware, engineering knowledge, or program access before using Claude. They divided work across sessions to conceal the full objective and attempted to bypass safeguards.
This qualification matters. The report does not show an uninformed person asking one question and receiving a complete weapon. It shows capable groups using AI to accelerate engineering, debugging, research, documentation, and procurement tasks.
Anthropic says it introduced classifiers that better detect high-yield explosives and weapons-development traffic. A classifier is an automated system that estimates whether a request belongs to a restricted category.
Such controls face an inherent problem. Many engineering tasks have legitimate civilian applications. Guidance, navigation, image processing, radio communications, materials analysis, and flight control can serve both commercial and military systems.
The report also describes biological research requests with dual-use characteristics. One case involved a funding proposal concerning changes to the chikungunya virus that might affect transmissibility and immune evasion.
Claude reportedly refused to help with portions of that work, and the user turned to another AI service. Anthropic says its older models were below the threshold for materially assisting a sophisticated biological researcher. It does not offer the same assurance for current systems.
Coverage from the biological safeguards case underscores that uncertainty. More capable models can support legitimate research, but the same competence complicates decisions about which requests to block.
Overblocking carries real costs. Scientists, public-health teams, and security researchers may need information that resembles restricted work. Underblocking can give a malicious actor assistance with experimental design or operational planning.
A model provider must make those decisions with incomplete context. A benign-looking request can be one fragment of a concealed program. A concerning request can be part of authorized defensive research.
The same weakness affects cyber safeguards. Attackers can divide objectives into ordinary tasks, switch accounts, use stolen credentials, or combine multiple providers. Open-weight models add another path without a central operator capable of closing an account.
Therefore, no provider can frame safety as a completed product feature. Safeguards create friction and improve detection, but they do not remove capability from the broader environment. The relevant measure is whether intervention raises attacker costs faster than capability lowers them.
On Anthropic’s AI Misuse Report supplies evidence of successful disruptions, but it cannot show the campaigns Anthropic never detected. It also cannot determine how many operators abandoned a project, migrated elsewhere, or continued with locally deployed systems.
That uncertainty should prevent two opposite mistakes. The cases do not prove that AI agents can independently execute every sophisticated operation. They also do not support dismissing the problem because humans still choose objectives.
Human direction and machine execution can coexist. In fact, that combination may be the most practical structure for harmful operations because it preserves judgment while scaling repeatable work.
Model Distillation Turns Customer Data Into a Security Question
The report’s final reversal is that AI providers are not only misuse platforms but also targets, data sources, and stolen computing resources.
Anthropic accuses several Chinese AI labs of running unauthorized distillation campaigns against Claude. Distillation uses outputs from one model to improve another model’s behavior or capabilities.
According to Anthropic, Alibaba generated more than 151 million exchanges between May and July 2026. Moonshot allegedly generated more than 23 million exchanges during the same period. DeepSeek produced more than 12.1 million exchanges across 14 days in July.
The company attributes more than 3.4 million exchanges to Zhipu over 17 days in June and July. It says Xiaomi generated more than 400,000 exchanges across 20 days in March and April.
These figures are Anthropic’s findings and have not been independently audited in the report. The named companies’ responses also matter when evaluating attribution, intent, and contractual claims.
Anthropic says the campaigns sought reasoning, coding, tool-use, and data-analysis capabilities. Operators allegedly used false identities, stolen cards, compromised API credentials, proxies, and thousands of accounts to evade access controls.
Some tested many prompt variations to extract hidden reasoning traces. One experiment reportedly sent over 12,000 different requests to identify successful extraction methods before launching a larger campaign.
The most concerning claim involves customer data. Anthropic says DeepSeek, Moonshot, and Xiaomi routed some user conversations to Claude for training purposes. Users allegedly were not told that another model provider would process their requests.
Anthropic reports that some exchanges contained names, email addresses, company information, developer credentials, and internal forecasts. Third-party routing services reportedly supplied additional conversations involving users in the United States and Europe.
These claims expand the meaning of AI supply-chain risk. A customer may think information is going to one application and model. In practice, requests can pass through routers, resellers, proxy services, evaluators, and hidden upstream providers.
Companies should ask vendors which models actually process submitted information. They also need clear answers about retention, training, regional routing, subprocessors, and access by human reviewers.
The episode creates a difficult symmetry. Anthropic criticizes other organizations for allegedly forwarding user data without consent. At the same time, its own report demonstrates how much model providers can infer through abuse monitoring.
Those situations are not equivalent, but they share a governance question. Sensitive information can travel farther than the user expects, whether through covert routing, security telemetry, or an investigation.
The answer cannot be a promise that confidential data will never move. Organizations need enforceable controls, verifiable routing, minimized retention, and logs that show which systems processed each request.
This is also why employees should avoid placing live secrets into unapproved AI tools. A polished interface does not establish where the request goes or how many intermediaries can access it.
What Defenders Should Watch Next
The next test is whether the security response changes the economics described in On Anthropic’s AI Misuse Report.
The first signal is cross-provider disruption. Anthropic says it shares findings with public and private partners when appropriate. The stronger test is whether one provider’s detection quickly disables related accounts, infrastructure, and stolen credentials elsewhere.
If cooperation improves, attackers will lose some ability to migrate intact workflows between services. If campaigns repeatedly reappear through fresh accounts and alternative models, account bans will remain temporary friction.
The second signal is evidence about agent autonomy. Future reports should separate AI-generated suggestions from actions executed through tools. They should identify where humans approved commands, corrected failures, selected targets, or interpreted ambiguous results.
That distinction will show whether agents are becoming more independent or simply making skilled operators more productive. Both outcomes matter, but they require different defenses and policy responses.
The third signal is transparency about investigation and privacy. Providers should explain what data they retain, how abuse classifiers trigger review, and how they limit investigator access. They should also report false positives and appeals where disclosure does not aid attackers.
For security leaders, the immediate response is practical. Inventory AI credentials, review model routers, reduce long-lived tokens, and monitor unusual automated activity. Test whether compromised developer devices can expose model sessions or related production secrets.
Teams should also revise incident-response plans for machine-speed iteration. A blocked domain or detected payload may prompt an automated revision within minutes. Defenders need coordinated identity revocation, endpoint containment, cloud logging, and vendor communication.
Knowledge workers should verify where sensitive prompts travel. Before submitting source code, internal documents, credentials, research data, or customer records, confirm the approved provider and its processing terms.
Anthropic’s report does not establish an autonomous cyber apocalypse. It documents something more immediate: humans using AI to turn expertise into repeatable infrastructure across cybercrime, surveillance, propaganda, and weapons work.
The question for the next report is not whether misuse continues. It is whether defenders force attackers to spend more identities, money, time, and expertise for every successful operation. That measurable contest will reveal whether safeguards are containing the shift or merely documenting it.



