Microsoft Copilot Cowork Sandbox Bypass Turned a Trusted Skill Into an Exfiltration Channel
Microsoft mitigated a Copilot Cowork vulnerability after researchers showed that one malicious Skill could bypass the product’s sandbox and exfiltrate work data. The Microsoft Copilot Cowork sandbox bypass reportedly created a command channel between an attacker’s server and the agent’s isolated environment.
PromptArmor says that channel could reach data available through Outlook, SharePoint, Teams, connected plugins, and the active session. The researcher reported the issue on June 24, 2026, and Microsoft confirmed mitigation on August 19.
The finding challenges a central promise behind workplace agents. A sandbox can isolate code, but isolation means little when a trusted service becomes an unmonitored route across its boundary. Microsoft says Cowork now evaluates Skills and warns users to upload them only from trusted sources.
The reported flaw is no longer an open zero-day, according to the disclosure timeline. However, the design lesson extends beyond one fixed implementation. Enterprise agents combine untrusted instructions, executable assets, organizational data, and authenticated tools inside one workflow.
That combination makes the trust boundary harder to define than it was for conventional business software. The central question is no longer whether an agent runs inside a sandbox. Buyers must ask which services cross that sandbox, what those services accept, and what activity security teams can observe.
How the Copilot Cowork File Exfiltration Chain Worked
The exploit reportedly turned a legitimate file-transfer service into a two-way command channel that the sandbox’s network restrictions did not stop.
Copilot Cowork is an agentic Microsoft 365 system that can plan and execute multi-step tasks. It can work with documents, search organizational information, create files, send messages, and call specialized Skills.
A Skill is a reusable instruction package that guides the agent through a particular workflow. Microsoft supports built-in, custom, shared, and plugin-based Skills, depending on the Cowork environment and administrative configuration.
PromptArmor’s example began with a normal business task. A user asked Cowork to compare a contract with a proposal through a document-consistency Skill downloaded from an outside source.
The Skill produced the requested comparison, so the visible task appeared successful. However, PromptArmor says a bundled script also invoked a file synchronization service operating outside the agent’s sandbox.
Cowork reportedly used that service to move files between external storage and the isolated working environment. The service accepted a URL identifying the file it needed to retrieve.
According to the researcher’s sandbox bypass analysis, malicious code could supply an attacker-controlled URL instead. That behavior let the script cause an outside service to contact infrastructure that the sandbox could not reach directly.
The attacker’s server returned a file containing a command. The malicious script read that file, executed the command inside Cowork, and encoded the result into another requested URL.
That second request carried the command output back to the attacker’s server. Repeating the process every few seconds reportedly established a command-and-control loop, meaning an attacker could issue new instructions based on earlier results.
This was more than a single outbound request. It created an interactive pathway into an environment connected to organizational services.
PromptArmor says its demonstration used Cowork’s Model Context Protocol server. MCP is a standard interface through which an AI application can call connected tools and retrieve data.
The researcher claims the attacker could use commands to query services available through that MCP connection. The demonstration included listing Outlook messages and retrieving the contents of an email thread.
The same access path reportedly exposed SharePoint files, session history, plugin data, and other information available to the active user. The actual scope would depend on that user’s permissions and connected services.
One detail made the reported behavior especially concerning. PromptArmor says selecting the stop control did not terminate processes already running in the background.
The visible agent turn could end while the malicious script continued polling the attacker’s server. A user might therefore believe the task had stopped while the covert channel remained active.
PromptArmor reported the issue to Microsoft on June 24. Microsoft requested additional information on July 24, discussed a fix through early August, and confirmed mitigation on August 19.
The public evidence comes primarily from PromptArmor’s technical account and demonstration. Microsoft has not published a detailed advisory explaining the code change, affected versions, or telemetry available for retrospective investigation.
Why the Microsoft Copilot Cowork Sandbox Bypass Matters
The vulnerability attacked the service connecting the sandbox to useful data, rather than defeating isolation through a conventional escape.
A traditional sandbox tries to contain untrusted code by restricting files, processes, devices, and network connections. That model works only when every route crossing the boundary applies equally strict validation.
Modern agents complicate this arrangement because useful work requires controlled exceptions. An agent must receive documents, return generated files, call tools, access enterprise systems, and preserve enough state to finish long tasks.
Each exception becomes a broker between the isolated environment and something more trusted. The broker can be safe when it validates both destinations and data flows. It becomes dangerous when untrusted code can repurpose it.
PromptArmor’s account does not describe a classic memory-corruption exploit or operating-system escape. The malicious Copilot Cowork Skill allegedly remained inside the sandbox while abusing a privileged service outside it.
That distinction matters for enterprise security reviews. A vendor can truthfully say that code runs in isolation while overlooking a broker that makes arbitrary external requests for that code.
The architectural tension sits at the center of Cowork’s value. Microsoft presents the product as an agent that moves beyond chat and completes work across Microsoft 365.
In Microsoft’s Cowork product announcement, the company highlighted inbox workflows, research, document generation, integrations, and reusable Skills. Those capabilities require access to valuable business context.
Cowork’s current documentation says tasks process user files inside a temporary, isolated environment within the Microsoft 365 service boundary. It also says the environment is removed after the task finishes.
That security model limits direct exposure, but it does not eliminate risk from connected tools. An isolated environment can still become a launch point when a trusted intermediary accepts attacker-controlled input.
This is why the Microsoft Copilot Cowork sandbox bypass pressures both Microsoft and enterprise buyers. Microsoft must show that the repaired boundary covers every broker, not only the synchronization route identified by one researcher.
Customers must also revisit assumptions about user permissions. Cowork runs with the active user’s access, so an exploited task does not need a separate account compromise to reach permitted data.
Least privilege still reduces the blast radius. It does not prevent misuse of access that the user legitimately holds.
An employee comparing two documents might have permission to read sensitive mail, deal folders, or internal discussions. Those permissions can become available to an agent through its approved connections.
The malicious input also arrived as a reusable workflow component, not an obviously hostile executable. That packaging lowers user suspicion because Skills are designed to look like productivity extensions.
The exploit therefore joins two security problems that organizations often manage separately. One is software supply-chain risk from downloaded packages. The other is authorization risk from agents acting as authenticated users.
Security programs need to handle them together. Reviewing code without mapping accessible data misses the impact, while governing permissions without inspecting Skill assets misses the entry point.
A Malicious Copilot Cowork Skill Can Still Look Useful
The most dangerous Skill is not one that visibly fails, but one that completes its assigned job while performing a hidden second task.
PromptArmor’s demonstration used a document-comparison workflow because it reflects an ordinary knowledge-work request. The Skill reportedly generated a complete consistency report while the malicious script opened the covert channel.
That dual behavior weakens a familiar safety signal. Users often treat a correct output as evidence that the tool operated as intended.
For an agent, output quality and execution integrity are separate questions. A useful report does not reveal every script, service call, or data request made while producing it.
Microsoft’s current custom Skill guidance says Cowork evaluates Skills automatically. The documented checks vary by reach and risk level.
Static checks examine structure, bundled code, and text for prompt-injection patterns. Behavioral checks evaluate outputs, actions, tool use, conflicts, and performance under realistic prompts.
The documentation describes additional gates for higher-risk deployments. These can include asset validation, trust and safety tests, adversarial evaluation, regression testing, and human review.
Microsoft also gives users a direct warning: only upload Skills from sources they trust. That advice acknowledges that automated inspection cannot turn arbitrary third-party code into a safe dependency.
The timing of those documented controls deserves careful treatment. Microsoft’s live documentation reflects the current product, not necessarily the exact configuration PromptArmor tested before August 19.
It is therefore unsafe to claim that every present control failed during the demonstration. It is equally unsafe to assume the listed checks would detect every variation of the attack.
Static scanning has inherent limits. Malicious behavior can be split across files, concealed behind normal functions, downloaded later, or triggered only under particular conditions.
Behavioral testing also samples a limited set of executions. A Skill can act safely during evaluation and activate harmful logic after a date, on a target tenant, or when specific data appears.
Recent academic work treats agent Skills as a software supply-chain surface rather than simple prompt templates. The SkillGate research evaluated a hybrid scanner against a benchmark containing 1,650 Skill packages.
Its authors reported an F1 score of 0.817 and a false-positive rate of 1.13 percent. Those results support runtime screening, but they also show that detection remains probabilistic.
The comparison is not a direct evaluation of Cowork, and the paper focuses on coding-agent Skills. Still, its threat model closely matches the broader problem.
A reusable instruction package can include scripts, trusted-looking documentation, and hidden behavior. Installing it expands the agent’s effective code and instruction base.
Traditional application stores address similar risk through signing, review, reputation, rapid removal, and permission declarations. Agent Skills require those measures plus visibility into model-directed execution.
The model can choose when and how to invoke supporting files. That flexibility makes a Skill more adaptable, but it also makes a complete behavior manifest difficult to produce.
Microsoft supports organizational sharing and App Store plugins alongside personal custom Skills. Administrators can govern plugin availability, deployment, connectors, and assigned users.
Those controls create a stronger distribution path than downloading an unknown archive. They do not remove the need to inspect privately shared or personally uploaded packages.
Enterprises should treat a malicious Copilot Cowork Skill like an untrusted application dependency. Provenance, versioning, approval, and revocation matter as much as the natural-language instructions it contains.
The Sandbox Promise Met the Reality of Connected Agents
The primary conflict is between containment as a security promise and connectivity as the feature that makes an enterprise agent valuable.
Microsoft says Cowork can send email, schedule meetings, create documents, post to Teams, search organizational information, and manage files. These actions transform it from a chatbot into an operational system.
Every added connector raises the value of a successful task. It also increases the potential impact when execution integrity fails.
The reported exploit did not need Cowork to receive more permissions than designed. It allegedly converted the existing authenticated tool layer into an attacker-controlled interface.
This is the reversal enterprise buyers should remember. A sandbox protected the execution environment, but an external synchronization service reportedly gave code a route around its network policy.
The same pattern can appear across agent platforms. Sandboxed agents commonly depend on browser automation, artifact stores, tool gateways, MCP servers, credential brokers, and connector runtimes.
Security teams often review each component independently. Attackers look for combinations.
A file service that seems low risk can become a network proxy. A tool endpoint intended for the agent can become a data-access API for malicious code.
A background worker designed for reliability can preserve an attack after the visible task stops. None of those components needs to look dangerous in isolation.
The immediate comparison is not Microsoft versus one rival. The more useful comparison is the industry’s containment promise against the operational reality of connected agents.
Anthropic, OpenAI, Microsoft, Google, and coding-agent vendors all face versions of this tension. Their products gain usefulness by reading more context and taking more actions.
The reported Cowork flaw is one implementation-specific example. It should not be generalized into evidence that every sandbox or MCP deployment has the same vulnerability.
However, it does show why a sandbox label cannot serve as a complete security assessment. Buyers need a data-flow model that includes every service with access across the boundary.
They also need clarity about process lifetime. Microsoft documents pause and cancel controls for Cowork tasks, but the PromptArmor test reportedly found that a background process survived the visible stop action.
Microsoft may have changed this behavior as part of mitigation, or it may have closed only the network route. The public disclosure does not explain the repair in that detail.
That verification gap matters. Organizations cannot determine from the public timeline whether Microsoft added destination validation, changed service authorization, terminated background jobs, improved detection, or combined several controls.
The absence of a detailed advisory does not mean the mitigation failed. It means customers must seek assurance through tenant guidance, support channels, audit data, and controlled testing.
Security architecture should assume that one preventive control will eventually miss a hostile package. A resilient design then limits what that package can reach and makes abnormal behavior visible.
For connected agents, that means restricting outbound destinations, authenticating broker requests, binding services to specific tasks, and separating read access from action permissions.
It also means revoking credentials when a task ends. A cancelled agent session should terminate related processes and invalidate any temporary authorization created for that execution.
Finally, monitoring must connect agent activity with conventional security telemetry. A Skill invocation, file transfer, MCP call, and unusual outbound request may look harmless in separate consoles.
Together, they can describe an attack chain.
Mitigation Does Not Close the Verification Gap
Microsoft has confirmed mitigation, but customers still lack enough public detail to reconstruct exposure or validate every affected control.
PromptArmor says Microsoft confirmed the issue was mitigated on August 19, nearly eight weeks after the initial disclosure. That timeline indicates coordinated remediation rather than an unresolved public exploit.
The researcher did not publish evidence of widespread exploitation. The demonstration proves a technical path under tested conditions, not the number of tenants affected in practice.
No public incident count, affected-version range, vulnerability identifier, or compromise indicator appears in the disclosure. Readers should not interpret the proof of concept as evidence of a broad breach.
The reverse conclusion would also be premature. Without a detailed Microsoft advisory, organizations cannot assume that absence of disclosed incidents means no malicious use occurred.
Retrospective investigation depends on telemetry. Administrators need to know whether Cowork logs expose uploaded Skill versions, script execution, broker requests, MCP calls, and background process lifetime.
Microsoft says Cowork activity can appear in unified audit logs and that Purview policies apply to the service. The current administrative documentation also describes controls for plugins, models, browser use, and automated tasks.
Those controls are relevant, but general audit coverage is not the same as detection coverage for this exploit. A log can record an allowed service request without marking its destination as hostile.
Organizations that tested Cowork before August 19 should ask Microsoft which events can identify the vulnerable behavior. They should also preserve relevant audit records before retention windows expire.
The most important review concerns Skill provenance. Teams should inventory custom Skills, uploaded archives, bundled scripts, and organization-shared packages used during the affected period.
Unknown or unverifiable packages deserve removal pending review. Security teams should compare cryptographic hashes where available, because a familiar Skill name does not establish file integrity.
Administrators should then map each Skill to the users who ran it and the data those users could access. This produces a more accurate exposure estimate than scanning Skill text alone.
Outbound network telemetry may provide another signal. Requests to unfamiliar domains, repeated polling intervals, or encoded data inside URL query strings deserve investigation.
However, the reported requests came through a service outside the sandbox. Endpoint monitoring on an employee’s device might therefore miss them.
Cloud-side logs and vendor telemetry become essential. Customers should ask whether Microsoft can expose the originating task, Skill, user, tenant, and requested destination for brokered transfers.
Organizations also need a policy for future uploads. Allowing any user to import a Skill from the public internet turns trust evaluation into an individual decision.
A safer model uses an internal registry with ownership, review status, approved versions, and expiration dates. High-risk Skills should receive both code review and runtime testing.
Shared Skills need change control after approval. A benign package can become dangerous through an update, compromised maintainer account, or swapped companion file.
Approval should therefore apply to a specific version, not a permanent name. Re-review should follow every material change.
These steps do not imply that Cowork is uniquely unsafe. They reflect the level of governance appropriate for any agent that can reach mail, documents, chats, and business applications.
Knowledge workers should also keep sensitive source material within clearly scoped repositories. Better knowledge management can reduce unnecessary data exposure when teams organize access around actual work needs.
The goal is not to remove useful context from every agent. It is to prevent one convenient workflow from inheriting an employee’s entire digital reach without deliberate review.
Three Signals Will Show Whether Agent Security Is Catching Up
The next test is whether Microsoft turns one mitigation into measurable, tenant-visible controls for every pathway crossing Cowork’s sandbox.
The first signal is a detailed Microsoft explanation of the fix. Customers need to know whether Cowork now validates synchronization destinations, binds requests to approved storage, and terminates processes when tasks stop.
A technical advisory would strengthen confidence because administrators could test the relevant boundaries. Silence would not prove continuing exposure, but it would keep verification dependent on private support channels.
The second signal is richer tenant telemetry. Security teams should watch for new Cowork events covering Skill hashes, bundled script execution, MCP operations, broker destinations, and task-linked background processes.
These records must be usable through existing detection systems. An activity log visible only inside one Cowork session will not support enterprise-scale threat hunting.
The third signal is stronger Skill governance. Microsoft’s documented evaluation system already describes static, behavioral, adversarial, regression, and human checks at different risk levels.
The key question is how consistently those controls apply to imported, personal, shared, and store-distributed Skills. Buyers should look for signed packages, fixed-version approvals, centralized allowlists, and rapid revocation.
Microsoft should also clarify whether a Skill’s companion files receive the same scrutiny as its main instruction file. The PromptArmor scenario depended on malicious bundled code, not only deceptive prose.
These signals will either strengthen or weaken the broader case for autonomous workplace agents. Better verification would show that vendors are treating Skills as executable supply-chain components.
Limited visibility would leave customers carrying a large share of the risk. They would be asked to trust a repaired sandbox without seeing the boundaries that changed.
For enterprises evaluating Cowork now, the practical response is measured deployment. Confirm the mitigation, restrict who can upload Skills, review existing packages, minimize access, and monitor connected services.
Do not treat a correct agent output as proof of safe execution. Do not treat a sandbox badge as proof that every external service obeys the same boundary.
The Microsoft Copilot Cowork sandbox bypass has reportedly been fixed, but its architectural warning remains. Agents concentrate instructions, code, credentials, and organizational context into one execution path.
Before expanding deployment, ask one concrete question: can your security team reconstruct every Skill, process, tool call, and external request behind a completed task? If the answer is unclear, make that visibility a requirement before granting broader access.



