top of page

Nation-State Hackers Use AI Across the Attack Chain

Aug 4
12 min read

Google has documented state-backed hackers using Gemini across nearly every attack stage, despite limited evidence of fully autonomous campaigns.

The finding changes the security debate. The immediate danger is not an independent AI hacker operating without supervision. It is a human operator who can research targets, write scripts, troubleshoot malware, and create convincing identities faster.

This google news story centers on research from Google Threat Intelligence Group, or GTIG. The group analyzes malicious activity using incident-response data, threat research, and visibility into attempted Gemini misuse.

Google found activity linked to China, North Korea, Iran, and Russia. These actors used AI for reconnaissance, vulnerability research, coding, social engineering, and information operations.

However, Google did not report that any state-backed group had automated an entire intrusion through Gemini. Human operators still selected targets, managed access, and made important operational decisions.

That distinction creates the article’s central tension. AI is already spreading across the attack chain, but broad adoption does not equal complete autonomy.

The development still puts defenders under pressure. Security teams must recognize familiar attacks that can now move faster, change more often, and appear more convincing.

Google News Reveals AI Across the Intrusion Cycle

The important change is AI’s movement from isolated experiments into repeatable tasks across the intrusion cycle.

Google published its detailed AI threat assessment in February 2026. The company said it had watched adversaries gradually integrate AI into their operations over several years.

That integration now covers many phases of an intrusion. Operators use models before an attack to research organizations, employees, technologies, and publicly known vulnerabilities.

During an operation, they can request code, debug scripts, translate messages, or develop new social-engineering approaches. After gaining access, they can seek help with commands, data processing, or additional tooling.

Google’s AI threat findings identified five broad categories of malicious activity. These included model extraction, AI-assisted operations, agentic experimentation, AI-integrated malware, and underground jailbreak services.

Model extraction uses repeated interactions to transfer useful behavior from a capable model into another system. This makes the model itself an intelligence and intellectual-property target.

AI-assisted operations were more immediately relevant to enterprise defenders. Google observed government-backed attackers using Gemini for coding, scripting, target research, vulnerability analysis, and post-compromise support.

Agentic AI describes systems that can plan steps, use tools, and adjust their actions toward a goal. Google saw attackers experimenting with these capabilities for reconnaissance and software development.

AI-integrated malware goes one step further. Instead of using a chatbot before deployment, the malicious program communicates with a model while it operates.

Google cited HONESTCUE, an experimental malware family that used the Gemini application programming interface. The malware sought generated code for downloading and executing another malicious component.

This does not mean HONESTCUE functioned as a self-directed digital operative. It demonstrates that attackers are exploring how a model can become part of malware’s operating logic.

The difference matters because an offline script behaves predictably after deployment. Model-connected malware can request new instructions or generate responses based on its environment.

That flexibility gives attackers more options. It also introduces dependencies, including internet access, usable accounts, reliable prompts, and responses that do not break the operation.

Google said it disables projects and accounts associated with detected malicious activity. That makes provider enforcement another obstacle for attackers using commercial models.

Underground services try to bypass those controls. Some claim to offer unrestricted systems while reportedly relying on jailbroken commercial interfaces or open-source components.

The resulting ecosystem resembles an access market. Attackers do not need to train a frontier model when they can obtain temporary, disguised, or stolen access.

This broad integration is why the headline matters. AI is no longer confined to better phishing messages or experimental malware samples.

It has become a general-purpose assistant that can support many points in a campaign. Each use can save time, reduce language barriers, or extend an operator’s technical reach.

State-Backed Groups Are Applying AI to Different Problems

Nation-state adoption is uneven because each group applies AI to its existing missions, resources, and operational habits.

Google linked attempted Gemini misuse to actors associated with several governments. The observed patterns reflected familiar national priorities rather than one shared AI playbook.

Some North Korean operators used AI to research job roles, compensation, and employees at cybersecurity or defense companies. That information can support fraudulent recruitment and insider-access schemes.

Another North Korean group reportedly consulted Gemini multiple days each week. Its operators sought technical assistance, troubleshooting support, and help generating malware code when their work stalled.

These examples show AI lowering practical barriers. An operator can obtain quick explanations without waiting for a specialist or searching through many technical sources.

North Korean campaigns already combine financial motives, espionage, and employment fraud. AI-generated personas and altered images can strengthen those operations without changing their basic purpose.

Palo Alto Networks documented suspected North Korean accounts using image tools to create fraudulent identities. Its incident response research also described fabricated companies supported by synthetic or repurposed profiles.

Iran-linked actors used AI in a different way. Google said one group significantly expanded reconnaissance against intended victims.

Reconnaissance is the process of collecting information before an intrusion. It can include mapping an organization’s employees, systems, suppliers, technologies, and exposed services.

Faster research helps an attacker personalize a lure or identify a promising access route. It can also connect scattered public information that a human might otherwise overlook.

Google observed China-linked group APT31 experimenting with agentic capabilities to automate reconnaissance and operate at greater scale. The group did not hand an entire campaign to a model.

Instead, the reported activity placed automation around a defined task. That approach is more practical because reconnaissance usually carries less operational risk than exploitation or lateral movement.

An incorrect research result wastes time. An incorrect command inside a victim’s network can expose the operator, crash a system, or destroy valuable access.

Russia, China, Iran, and North Korea also used Gemini for information operations, according to Google. Tasks included generating articles, personas, and other campaign assets.

Information operations aim to shape perceptions rather than compromise a computer directly. Generative models are well suited to producing variations of text and imagery at low marginal effort.

Yet output volume does not guarantee influence. A campaign still requires distribution, credible accounts, audience targeting, and narratives that resonate with real communities.

Google’s findings therefore describe augmentation, not uniform sophistication. A well-resourced espionage team and a fraudulent worker operation can use the same model for different goals.

The common benefit is convenience. Models combine translation, research, coding assistance, and content generation inside one interface.

That convenience reduces friction between attack stages. It does not erase the expertise needed to choose targets, preserve access, and avoid detection.

The Real Contest Is Human Direction Versus AI Automation

The defining conflict is not attacker AI against defender AI, but human-directed operations against increasingly autonomous workflows.

Google’s evidence shows models influencing almost every intrusion stage. It does not show state groups routinely delegating complete intrusions to autonomous systems.

John Hultquist, GTIG’s chief analyst, told cybersecurity reporting that countries were still determining where AI offered more benefit than friction.

That uncertainty explains why adoption appears task-specific. Attackers favor work that is repetitive, searchable, language-heavy, or easy to verify.

Reconnaissance fits that profile. So do code explanations, script generation, vulnerability summaries, translation, and the production of social-engineering content.

High-risk actions create a harder problem. Exploitation, persistence, credential use, and lateral movement depend on accurate context from a changing environment.

A model can produce a plausible command that is technically wrong. It can misunderstand a network, select a noisy technique, or expose its controller.

State-backed espionage groups value stealth because a durable foothold can remain useful for months. Faster automation provides little benefit if it reveals an operation immediately.

Cybercriminal groups may accept that tradeoff more readily. They often prioritize volume, speed, and short-term returns across many potential victims.

This difference suggests that AI’s effects will not arrive evenly. Smaller criminal operations can gain capabilities that previously required specialized staff.

Elite state groups can use the same tools to remove routine work. However, their experienced operators may remain cautious about autonomous execution.

The opposing routes are therefore clear. Human-directed augmentation offers control and lower operational risk, while autonomy offers scale and speed.

Most public evidence still favors augmentation. Humans define objectives, check outputs, and decide when to act.

However, the boundary is moving. AI-integrated malware and agentic reconnaissance allow models to perform more than advisory work.

HONESTCUE illustrates this transition. The malware incorporated model access into its execution chain, even though the broader operation still depended on conventional attacker infrastructure.

Google’s May 2026 research added stronger evidence. GTIG identified what it believed was the first threat actor using an AI-developed zero-day exploit.

A zero-day is a vulnerability unknown to the affected vendor when attackers prepare to exploit it. Defenders have no available patch at that moment.

Google said the criminal group intended to use the flaw in mass exploitation. The company notified the affected organization and law enforcement, potentially stopping the planned campaign.

The affected product and threat actors were not publicly named. That limits independent assessment of the exploit and its development process.

Still, the case pushes AI closer to a decisive attack stage. Vulnerability discovery and exploit creation have traditionally required specialized knowledge and substantial testing.

Google later described a broader move toward industrial-scale AI use in adversarial workflows. Its May threat tracker covered exploit generation, evasion, autonomous malware, information operations, and attacks on AI dependencies.

This progression does not establish a fully autonomous attack chain. It shows that separate AI-assisted components are becoming more capable and operationally relevant.

The central question is how quickly attackers can connect them. A reconnaissance agent, exploit generator, and adaptive malware component become more dangerous when coordinated reliably.

For now, human judgment remains the connective tissue. Defenders should not assume that this limitation will remain fixed.

AI Compresses Time Without Replacing Established Techniques

AI’s most credible near-term advantage is compressed execution time, not an entirely new class of cyberattack.

The observed campaigns still rely on recognizable methods. Attackers use stolen identities, malicious links, vulnerable software, remote-management tools, and compromised infrastructure.

AI assists those techniques by reducing the labor between steps. It can draft a lure, revise code, interpret an error, or summarize a target’s technology stack.

That acceleration matters because cybersecurity is already a race. Defenders need time to detect an intrusion, investigate its scope, and contain the attacker.

CrowdStrike’s 2026 threat report said AI-enabled adversary activity increased 89 percent year over year. It also reported a 29-minute average breakout time.

Breakout time measures the period between initial access and lateral movement into other systems. A shorter interval leaves defenders less time to isolate the first compromised account or device.

CrowdStrike tracked more than 280 named adversaries for the report. It described AI use across reconnaissance, credential theft, evasion, malware activity, and fraudulent personas.

The company linked Russia-nexus FANCY BEAR to LAMEHUG. That malware reportedly used a large language model to generate commands for reconnaissance and document collection.

Palo Alto Networks also found that evidence of large-scale nation-state AI adoption remained limited. It characterized 2025 as an early period of operational experimentation.

These findings are not contradictory. AI use can increase sharply from a low base while still remaining uneven across the threat landscape.

The available evidence supports a measured conclusion. AI is making existing work faster and more accessible before it consistently creates capabilities unavailable to skilled operators.

That distinction should shape security spending. Buying an AI-branded defensive product does not fix weak identity controls, unpatched systems, or excessive permissions.

Organizations still need reliable telemetry across endpoints, cloud systems, software pipelines, and identity services. AI-assisted attacks remain visible through the actions they perform.

A generated phishing message still needs delivery. Stolen credentials still produce authentication events. Malware still creates processes, accesses files, or communicates across a network.

However, defenders should expect greater variation. Attackers can rewrite scripts, personalize messages, and modify decoy content without the same manual cost.

Static indicators therefore lose value faster. A hash or exact text pattern may identify one artifact while missing many generated variants.

Behavioral detection becomes more important. It looks for suspicious actions and sequences rather than relying only on an artifact’s appearance.

Security teams also need faster triage. If AI helps attackers reduce the interval between research, access, and movement, slow alert review becomes a larger operational weakness.

Google is applying the same technologies defensively. The company has described Big Sleep for vulnerability discovery and CodeMender for automated code remediation.

This creates an asymmetric race. Attackers need one workable path, while defenders must protect many systems and handle false positives.

Defenders have their own advantage. They control enterprise identities, configurations, logs, and enforcement points when those systems are managed correctly.

The outcome will depend less on who possesses AI. It will depend on who connects automation to higher-quality context, permissions, and verification.

The Evidence Still Has Important Limits

Broad AI use is well supported, but claims of fully autonomous nation-state hacking remain ahead of the public evidence.

Google has unusual visibility into Gemini misuse. It can observe prompts, related accounts, and patterns that enterprise defenders cannot see.

That visibility creates valuable intelligence. It also means the public findings reflect one provider’s services and detection capabilities.

Attackers can use other commercial models, locally hosted systems, stolen accounts, or custom infrastructure. Google’s dataset cannot represent every adversarial AI workflow.

The reverse limitation also applies. Seeing a malicious prompt does not prove that an operator used the answer successfully inside a real intrusion.

Threat researchers strengthen assessments by connecting model activity with known actors, infrastructure, malware, or incident evidence. Public reports rarely disclose every supporting detail.

Providers must protect investigations, customers, and detection methods. That necessary secrecy prevents outsiders from reproducing some conclusions.

The February research also contained an explicit constraint. Google had not observed state groups using Gemini to automate large portions of a cyberattack.

The finding prevents an easy but misleading interpretation. Coverage across the attack chain is not the same as one continuous, automated attack chain.

A group can use AI for research on Monday, coding on Tuesday, and translation on Wednesday. Human operators may still connect every action.

Autonomous systems face reliability problems that become costly during intrusions. Models can hallucinate facts, mishandle credentials, repeat actions, or misunderstand tool output.

They can also create recognizable patterns. Excessive scanning, repetitive queries, or rapid command execution may make a campaign easier to detect.

State operators must consider attribution. An automated system that behaves unpredictably can expose infrastructure, reveal objectives, or create political consequences.

Commercial model controls add another constraint. Providers can identify abuse, disable accounts, change safeguards, and share indicators with defenders.

Attackers respond through jailbreaking, account cycling, middleware, and open-source models. Each workaround creates additional infrastructure that defenders can investigate.

The zero-day case also deserves caution. Google said code characteristics suggested AI involvement, including unusually explanatory comments and patterns associated with generated Python.

Those signals support an assessment, but they are not an infallible test. Humans can write similar code, and generated code can be edited.

Google shared limited information about the actor, target, and affected software. Independent researchers therefore cannot fully examine the evidence.

The reported zero-day case remains significant because Google connected AI assistance with a planned mass exploitation effort.

It should not be stretched into proof that models can independently discover, weaponize, and deploy unknown vulnerabilities against arbitrary targets.

Vendor metrics require similar care. CrowdStrike’s 89 percent increase reflects its definitions, observed customers, and tracked activity.

It does not mean 89 percent of all attacks use AI. Nor does it establish that AI caused every increase in speed or scale.

Security leaders should separate three claims. Attackers use AI broadly, some AI-assisted techniques work in real operations, and autonomous campaigns can replace expert teams.

Evidence strongly supports the first claim. It increasingly supports the second. The third remains unproven at broad operational scale.

That hierarchy keeps defensive planning grounded. Organizations should address demonstrated acceleration while preparing for greater autonomy without pretending it has already arrived everywhere.

What Security Teams Should Watch Next

The next phase will be defined by verified exploit development, adaptive malware, and measurable reductions in attacker operating time.

The first signal is another independently documented AI-assisted zero-day. Repetition would show that the May case was not an isolated event.

Researchers should look for evidence connecting model interaction to vulnerability discovery, exploit construction, testing, and operational deployment.

Clear technical artifacts would strengthen the assessment. Reproducible analysis matters more than broad claims that code merely appears machine-generated.

Several confirmed cases would increase pressure on software vendors. They would need shorter patch cycles, better exploit detection, and automated review of exposed code.

A lack of further cases would not prove the risk disappeared. It would suggest that reliable zero-day development remains difficult or poorly observable.

The second signal is malware that adapts during a real intrusion. PROMPTSPY, HONESTCUE, and LAMEHUG point toward model-connected execution, but autonomy exists on a spectrum.

The important test is whether malware interprets system state and chooses effective actions without constant operator guidance.

Defenders should watch for programs that generate commands, change collection priorities, or select evasion methods based on local conditions.

Successful adaptive behavior would strengthen the autonomy argument. Repeated operational failures would support continued human supervision.

This signal also affects network controls. Organizations may need to distinguish legitimate model access from compromised applications calling external or locally hosted systems.

The third signal is a measurable decline in attacker dwell and breakout times linked to AI-assisted workflows.

Faster movement would show that automation is improving campaign execution rather than only content production or research.

Security providers should explain their measurement methods. They should separate criminal activity from state-backed operations and identify where AI influenced the timeline.

If state groups preserve stealth while accelerating operations, defenders face the hardest scenario. They would lose response time without gaining obvious detection signals.

If automation creates noisy behavior, defensive AI may offset some speed advantages. Automated correlation and containment can act before a human analyst finishes an investigation.

Organizations do not need to wait for those signals before acting. They can reduce exposure by tightening identities, limiting service-account permissions, and monitoring unusual model access.

They should log AI interactions in sensitive engineering and security workflows. Model credentials deserve the same protection as other privileged application secrets.

Software teams should review third-party AI connectors and dependencies. Google has warned that compromised AI components can expose API credentials and provide routes into broader environments.

Recruitment and contractor processes also need stronger verification. Synthetic identities and AI-assisted interviews make informal checks less reliable.

The response should remain proportionate. Blocking every AI service can push usage into unmanaged accounts and remove visibility.

A better approach defines approved services, protects credentials, records access, and limits what automated tools can reach.

Security teams should also preserve human review for consequential actions. An agent should not receive broad production permissions because it performs low-risk analysis well.

The latest google news coverage is a warning about integration rather than science fiction. Attackers are attaching models to the work they already understand.

Defenders should ask one practical question: which slow, repetitive step in an intrusion would become most dangerous if an adversary automated it tomorrow?

That answer should guide the next control, detection, or exercise. The organizations that test those assumptions now will be better prepared when AI crosses from assistance into dependable autonomy.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page