Cohesity Agent Resilience Targets the Recovery Gap AI Agents Create
Cohesity introduced Cohesity Agent Resilience on September 16, extending recovery protection to AI agents and the systems they can change. The new capability arrives with a harder promise: automate more of cyber response without removing human approval.
That distinction matters because enterprise agents are becoming active operators, not passive assistants. They retain memory, call applications, update databases, and execute workflows. Detection can expose a harmful action, but detection alone cannot restore the agent or reverse every affected resource.
Cohesity is treating the agent, its state, and its connected infrastructure as one recoverable system. Commvault, Druva, and Rubrik are pursuing related approaches, turning AI agent recovery into a new contest across the data protection market.
Cohesity Agent Resilience Protects More Than the Agent
Cohesity Agent Resilience treats an AI agent's memory and configuration as recoverable operational state.
Cohesity announced the capability during its Catalyst 2026 virtual event, held across multiple regions on September 16 and 17. The company had previously said that Catalyst would introduce new controls for multi-agent environments and autonomous recovery workflows.
An agent's state includes the stored context and configuration that influence its behavior. If malicious instructions, configuration drift, or corrupted memory alter that state, restoring a database will not necessarily restore the agent's judgment.
Cohesity says its new capability creates point-in-time recovery options for that state. Administrators can return an affected agent to a known-good version instead of rebuilding it and losing its accumulated context.
The protection also extends to the infrastructure around the agent. According to the company's agent recovery overview, that scope can include applications, databases, file systems, memory stores, and other services the agent uses or manages.
That wider scope addresses a basic problem with agentic software. An agent can remain functional while its memory is corrupted, or it can recover while leaving harmful changes behind in connected systems.
Cohesity therefore separates two recovery tasks. One restores the agent's trusted state. The other identifies and recovers resources affected by its actions.
The company is using established data protection mechanisms for both tasks. These include snapshots, immutable backups, and isolated recovery environments, often called clean rooms.
A clean room is a separated environment used to inspect and restore systems without immediately reconnecting them to production. It gives responders space to check whether recovered components remain compromised.
The company also describes an agent topology that maps each agent to its memory, applications, databases, and supporting infrastructure. That relationship map is intended to show protection gaps and identify everything required for a trusted recovery.
Cohesity Agent Resilience initially supports Amazon Bedrock AgentCore and Bedrock Agents. Select customers can access it now, while general availability is targeted for the end of 2026. Microsoft Azure and Google platforms remain on the roadmap.
Those limits make this a focused first release rather than a universal agent recovery layer. Cohesity must prove that the model works across different clouds, agent frameworks, memory systems, and permission structures.
The announcement still changes the recovery boundary. Backup platforms traditionally protect data and applications. Cohesity now argues that an agent's operational memory belongs inside that boundary because it can directly shape production activity.
The reported launch details and roadmap are also documented in the original agent resilience coverage. The central idea is straightforward: enterprises need a way to recover the actor as well as the assets it touched.
AI Agents Put Traditional Recovery Plans Under Pressure
An AI agent expands the incident surface because one compromised decision can propagate across several connected systems.
Traditional recovery planning generally assumes that teams can identify damaged workloads, restore clean copies, and validate the resulting environment. AI agents complicate that sequence because they hold context and take actions across application boundaries.
Consider an internal support agent with access to tickets, customer records, and knowledge repositories. A corrupted instruction could cause it to disclose information, change classifications, or write incorrect material into several systems.
Restoring the ticketing database might recover deleted records. It would not remove the compromised instruction from the agent's memory or show which downstream actions require reversal.
The same problem appears in software operations. An agent that can modify configuration, create deployment requests, or manage cloud resources may produce a chain of validly authenticated but harmful changes.
Those actions might not resemble a conventional intrusion. The agent may use legitimate credentials and approved interfaces while operating from poisoned context or a faulty configuration.
This is why Cohesity frames AI agent recovery as a resilience problem, not only a monitoring problem. Monitoring records behavior. Recovery establishes a trusted point and restores the systems damaged after that point.
Vasu Murthy, Cohesity's chief product officer, summarized the gap directly: detection can show that an agent went off course, but it cannot undo the changes. That claim explains why the company protects both state and dependencies.
The pressure falls first on security and infrastructure teams. They must decide which agent memories matter, how frequently to capture them, and how to coordinate recovery across connected resources.
AI engineering teams face a related burden. They need to expose enough information for protection systems to identify agent versions, configurations, memory stores, permissions, and external dependencies.
Application owners also enter the recovery chain. A clean agent restore has limited value if its databases, identity controls, or business applications remain in an inconsistent state.
The organizational problem can become harder than the storage problem. Different teams may own the agent, its model access, the underlying data, and each connected application.
Cohesity's topology concept attempts to make those relationships visible before an incident. A current dependency map can tell responders which systems need investigation and which recovery order preserves consistency.
That mapping also supports recovery time objectives and recovery point objectives. An RTO defines how quickly a service should return, while an RPO defines how much recent state an organization can afford to lose.
Those measurements become less clear for an agent. Restoring yesterday's memory may erase useful context, while restoring today's memory may preserve the corruption that caused the incident.
Teams will therefore need policies that distinguish durable knowledge from transient state. They must also decide which agent actions can be reversed automatically and which require business review.
Cohesity's proposition creates pressure for other backup vendors because customers will expect recovery plans to follow this broader system boundary. Protecting only files or workloads looks incomplete once autonomous software can modify both.
The issue also reaches enterprises that have not adopted highly autonomous agents. Even limited agents can write summaries, classify records, call tools, or trigger workflows that influence later decisions.
Agent resilience is therefore not reserved for fully autonomous systems. The relevant threshold is whether the software can preserve state or make consequential changes without reviewing every action with a human.
The New Contest Is Cohesity AI Agent Recovery Versus Fragmented Controls
The market contest is not simply Cohesity against another vendor; it is unified AI agent recovery against disconnected monitoring, backup, and application controls.
Enterprises already have tools that address pieces of agent risk. Observability systems capture traces, identity platforms manage access, application logs record changes, and backup products preserve data.
The recovery problem appears between those layers. An incident team may identify a suspicious agent session but still lack a coordinated way to restore its memory and every affected dependency.
Cohesity wants its data platform to become that coordination layer. It can use its existing snapshots and immutable copies while adding knowledge about agent topology and state.
That approach gives an established backup provider a natural entry point. The company already manages recovery copies and isolated environments for many customers.
However, existing infrastructure does not automatically solve agent semantics. A platform must know which memory items, configuration files, prompts, tools, credentials, and application changes belong to a specific recovery point.
Competitors are moving toward the same problem from different directions. Rubrik's AgentCloud emphasizes agent operations, governance, observability, and a rewind capability for unintended actions.
Rubrik is also opening its cyber resilience data to customer agents through Model Context Protocol, or MCP. MCP is a standard interface through which AI systems can access external tools and data under defined controls.
That strategy lets agents reason over recovery information and initiate governed workflows. Recent MCP integration details show how quickly recovery platforms are becoming callable components inside broader AI systems.
Commvault has announced AI Protect, which is designed to discover agents and inventory their dependencies across connected environments. Its planned records include models, configurations, data sources, applications, and infrastructure.
Druva has described protection for agents, access to backup information through agents, and automated responses to suspected AI attacks. That combines agent recovery with agent-assisted investigation.
These approaches overlap, but each emphasizes a different control point. Cohesity centers recoverable state and connected resources. Commvault stresses recurring discovery. Druva combines protection with security investigation, while Rubrik links governance and reversible operations.
Customers should examine how deeply each platform understands agent dependencies. A product that captures configuration without mapping downstream resources solves only half the problem.
They should also examine framework coverage. Cohesity's opening support for AWS Bedrock provides a defined integration surface, but many enterprises build agents with custom frameworks and mixed cloud services.
A recovery platform must handle those environments without forcing every agent into one vendor's orchestration model. Open interfaces can help, but they also expand the permissions and security decisions that administrators must manage.
The competitive pressure is likely to favor vendors that connect three capabilities. They must discover agent relationships, preserve trustworthy versions, and orchestrate consistent recovery across dependent systems.
No vendor has yet shown that this can work universally across enterprise agent stacks. Product announcements establish direction, while production evidence will determine whether unified recovery beats fragmented controls.
Cohesity's advantage is its existing recovery foundation. Its challenge is proving that the foundation can represent the constantly changing state of agents with enough precision to support safe restoration.
Autonomous Cyber Resilience Still Depends on Human Judgment
Cohesity's automation plan can reduce repetitive work, but it cannot remove the need to decide what a safe recovery should preserve.
Alongside Cohesity Agent Resilience, the company outlined a broader vision called Autonomous Cyber Resilience. It uses agentic workflows to automate parts of a five-step resilience framework.
Those steps cover protection, recoverability, threat remediation, recovery practice, and continued improvement of data and AI risk posture. The proposed system would continuously discover assets, assess protection, validate recovery readiness, and update plans.
Cohesity describes a workflow in which administrators set business objectives through Cohesity Copilot. The platform then evaluates relevant workloads and recommends protection policies, threat scanning, and recovery rehearsals.
Humans would approve those recommendations before execution. During an incident, administrators could start automated workflows that assess impact, find indicators of attacker activity, and prepare an isolated recovery environment.
This model is more cautious than fully autonomous remediation. It gives software responsibility for discovery, analysis, preparation, and orchestration while keeping approval with human operators.
That boundary is important because recovery decisions involve conflicting objectives. The fastest available restore point might preserve corrupted state. An older restore point might remove the threat but lose recent transactions.
Agentic workflows can assemble evidence and test options, but organizations still need accountable people to choose between those outcomes. The correct choice depends on business priorities and incident context.
Cohesity says the model builds on RecoveryAgent, which already coordinates parts of incident response and recovery. It also plans to expand Cohesity Maestro, an interface layer connecting resilience capabilities to outside AI tools.
The company expects Maestro to work with systems including Claude, ChatGPT, Gemini, and its Helios management console. Cohesity has already described how Claude workflows can access its resilience intelligence through MCP and agent skills.
These integrations create useful flexibility. Security teams can reach recovery information from the AI environment where they already investigate incidents.
They also create another control surface. Any agent that can query recovery data or prepare actions needs tightly scoped permissions, reliable identity, comprehensive logging, and reviewable outputs.
A compromised automation system must not gain unrestricted access to both production assets and their recovery copies. Separation of duties remains important even when an agent coordinates the process.
The phrase autonomous cyber resilience can also hide different levels of automation. Automatic asset discovery carries less operational risk than automatic restoration of production applications.
Policy recommendations sit somewhere between those points. They can save administrative work, yet a poor recommendation becomes consequential when teams approve it without understanding its assumptions.
Cohesity's current automation tutorial shows that foundational capabilities already connect data discovery with ongoing protection decisions. Broader commercial automation will arrive incrementally.
That staged approach is sensible because the system needs evidence at each level. Discovery accuracy, policy quality, clean-room preparation, and recovery consistency should be measured separately.
The largest uncertainty is not whether agents can execute recovery steps. Automation software has coordinated infrastructure tasks for years.
The uncertainty is whether the platform can maintain a sufficiently accurate model of business dependencies while applications and agents change. A stale topology could produce a technically successful but operationally incomplete recovery.
Human approval does not eliminate that risk. Reviewers need clear explanations of what the workflow found, what it excluded, which recovery points it selected, and how confident it is.
Autonomous cyber resilience will be credible when it makes those judgments inspectable. Speed matters during an incident, but unexplained speed can multiply damage.
What Cohesity Agent Resilience Has Not Proven Yet
The announcement defines a credible recovery model, but production coverage and measurable recovery results remain unverified.
The first limitation is availability. Select customers can use Cohesity Agent Resilience now, while broad availability is targeted for the end of 2026.
A controlled rollout can help Cohesity refine agent discovery and recovery. It also means most prospective customers cannot yet compare the product against their full production environments.
The second limitation is platform scope. Initial support focuses on Amazon Bedrock AgentCore and Bedrock Agents, with Azure and Google environments still planned.
Enterprise agents often span several services. An agent may run in one cloud, retrieve documents from another platform, call a SaaS application, and store memory in an external database.
Recovering only the supported portion could create inconsistent state. Cohesity will need to show how topology and restoration work when part of that chain sits outside its direct control.
The third limitation involves recovery granularity. An agent's state can include prompts, short-term memory, long-term memory, tool definitions, model settings, access policies, and external records.
Not every component should return to the same timestamp. Some records may be valid after an incident began, while one poisoned memory item may be the actual cause.
Restoring an entire state bundle might remove legitimate work. Restoring too narrowly might leave the compromise intact.
Cohesity has not publicly supplied detailed performance results for these scenarios. Buyers should seek evidence about discovery accuracy, restore time, application consistency, and the amount of manual reconciliation required.
They should also ask how the system handles shared dependencies. Two agents may write to the same database or use the same memory store, making individual recovery difficult.
Identity introduces another unresolved issue. Returning an agent to a clean configuration does not necessarily revoke a compromised token or correct overly broad privileges.
A complete recovery process must coordinate with identity and access systems. Otherwise, the restored agent can inherit the same path that enabled the problem.
The fourth limitation concerns automated recommendations. Cohesity says its platform can assess posture and propose policies, scanning strategies, and recovery rehearsal plans.
Those outputs are company claims until customers validate them across real incidents and complex application estates. Buyers should evaluate recommendation quality rather than treating automation as an outcome by itself.
Cohesity's own research adds urgency but does not validate the product. Its fifth annual cyber resilience report found that 78 percent of surveyed organizations focus recovery efforts on restoring systems rather than maintaining business operations.
That figure supports the company's argument that technical restoration can fall short. It does not show that Agent Resilience or autonomous workflows close the gap.
The distinction between system recovery and business recovery is still useful. A restored application may depend on identity services, current data, employee access, and other applications before operations can resume.
Cohesity highlighted this issue during its Catalyst event, where the company framed recovery around restoring a minimum viable business rather than isolated infrastructure.
That framing raises the right standard. Customers should judge agent recovery by whether trusted business processes resume, not by whether a snapshot successfully mounted.
The fifth limitation is accountability. If an automated workflow recommends the wrong restore sequence, organizations need a record of why it made that choice and who approved it.
Governance cannot stop at a human approval button. Reviewers need enough context to make approval meaningful, especially when incident pressure encourages rapid action.
None of these questions invalidate Cohesity's direction. They define the evidence required to move from a plausible architecture to a dependable operational system.
Three Signals Will Show Whether the Strategy Works
Cohesity's strategy should be judged through platform coverage, verified recovery outcomes, and the boundaries placed around automation.
The first signal is delivery of general availability by the end of 2026. That release should include clear documentation for protected state, dependency discovery, restore sequencing, and unsupported configurations.
Broader cloud support will matter just as much as the date. Progress on Microsoft and Google integrations would show that Cohesity can extend beyond a tightly controlled AWS implementation.
A delayed release or limited framework coverage would weaken the claim that the platform can become a general recovery layer for enterprise agents.
The second signal is customer evidence. Cohesity needs examples that show an agent returning to trusted state while connected applications and data remain consistent.
Useful evidence would report recovery time, the number of dependencies discovered, failed or incomplete restores, and the manual work required afterward.
The strongest proof would involve realistic incidents rather than prepared demonstrations. Configuration corruption, poisoned memory, excessive permissions, and unintended downstream writes should produce different recovery challenges.
Customers should also look for repeated exercises. A successful one-time recovery says less than regular tests showing that the system stays current as agents and applications change.
The third signal is how Cohesity expands autonomous cyber resilience. The company says it will add automation as the model approaches commercial availability.
The critical question is which decisions remain recommendations and which become executable actions. Automatic discovery and clean-room preparation are different from automatically selecting production restore points.
Transparent approval controls would strengthen Cohesity's case. Those controls should show the proposed action, supporting evidence, expected impact, excluded assets, and a rollback path.
The strategy would weaken if autonomy advances faster than auditability. Recovery automation should reduce response time without obscuring responsibility.
Competitor behavior will provide additional context around all three signals. Commvault, Druva, and Rubrik are unlikely to leave agent protection as a narrow feature category.
Their responses can establish common expectations for topology mapping, agent rewind, immutable state, and open integrations. They can also expose gaps in Cohesity's coverage.
For enterprise buyers, the practical step is to inventory what their agents can change now. Teams should identify memory stores, credentials, applications, databases, and infrastructure attached to each important agent.
That exercise will reveal whether existing backup plans cover the full operational chain. It will also show where an AI agent recovery product must integrate with identity, application, and security controls.
Organizations building internal knowledge workflows should apply the same discipline to information dependencies. A maintained AI knowledge base can help teams understand which source material guides human and automated decisions.
Cohesity Agent Resilience points toward a necessary change: recovery planning must account for software actors that remember and act. The product's success now depends on whether it can restore those actors without losing consistency, context, or control.
Before trusting automation with recovery, ask one concrete question: can the system explain exactly what it will restore, what it will leave untouched, and why? If the answer is incomplete, keep the human approval gate firmly in place.



