top of page

GitLab AI Gateway Vulnerability Breaks the Sandbox and Enables Command Execution

4 days ago
12 min read

GitLab has patched a GitLab AI Gateway vulnerability rated 9.9 out of 10 after researchers found a path to arbitrary command execution. The flaw affects self-hosted gateways and requires an authenticated user with access to the GitLab Duo Agent Platform. A specially crafted custom flow can escape the prompt template sandbox and run commands with the gateway process’s privileges.

The incident is narrower than a remotely exploitable flaw affecting every GitLab deployment. GitLab-hosted gateways were patched by the company, and customers using those services do not need to update the gateway themselves. Organizations operating a self-hosted AI Gateway carry the immediate responsibility.

That distinction creates the central tension. Self-hosting gives an organization greater control over AI processing, infrastructure, and data pathways. It also transfers patching, monitoring, and containment duties to the customer. In this case, the component placed inside the trusted environment became a command-execution target.

The vulnerability is tracked as CVE-2026-90970. GitLab released AI Gateway versions 19.2.4, 19.3.2, and 19.4.1 on October 2, 2026. The company urged affected operators to update immediately, but its public advisory did not identify a workaround or provide detailed compromise-detection guidance.

The GitLab AI Gateway Vulnerability Requires an Immediate Update

The urgent fact is simple: affected self-hosted gateways need a patched AI Gateway version, not merely an updated GitLab application.

GitLab’s critical patch advisory identifies three fixed versions: 19.2.4, 19.3.2, and 19.4.1. The relevant version number belongs to the AI Gateway component. It should not be confused with the version of a connected GitLab instance.

The affected range starts with AI Gateway 18.1.6. It includes versions before 19.2.4, version 19.3 before 19.3.2, and version 19.4 before 19.4.1. Organizations running an older supported branch must select the appropriate patched release.

The wording also matters for installations on the 18.x and 19.1 lines. GitLab did not list a fixed gateway release below 19.2.4 in the October advisory. Operators on those lines should not assume that remaining on an older branch provides protection.

A self-hosted AI Gateway runs as a separate service, typically through a container or Helm deployment. Updating the main GitLab application does not automatically prove that the gateway image was replaced. Administrators must inspect the gateway deployment itself and confirm its running image tag.

GitLab said it contacted self-hosted gateway customers before publishing the advisory. It also said that a fix had already reached GitLab-hosted gateways. That protects GitLab.com, GitLab Dedicated, and self-managed instances connected to a GitLab-hosted gateway.

This scope makes asset identification the first operational challenge. A security team might know that its organization uses GitLab Duo without knowing which gateway model supports it. The service can be operated by GitLab or deployed within the customer’s environment.

Teams should verify that distinction through deployment records, container inventories, Helm releases, and GitLab Duo configuration. They should avoid inferring gateway ownership from the GitLab product name alone. A self-managed GitLab instance can still use a GitLab-hosted gateway.

The CVE record assigns the vulnerability the maximum confidentiality, integrity, and availability impacts within its CVSS 3.1 vector. The score is 9.9 rather than 10 because exploitation requires low privileges. It does not require user interaction, and the vulnerable service is reachable over a network.

Low privilege does not mean anonymous access. GitLab says an attacker must be authenticated and have Duo Agent Platform access. That requirement narrows the initial attacker population, but it includes compromised accounts and potentially malicious insiders.

The flaw crosses a security boundary after that initial access. A user authorized to work with AI flows should not inherit permission to execute operating-system commands on the gateway. The vulnerability turns application-level access into control over a more sensitive execution environment.

Operators should treat the update as an independent change with independent evidence. A completed change ticket for GitLab does not establish that the AI Gateway is safe. The correct evidence is the running gateway version and a verified rollout across every replica.

That verification should include development, staging, disaster-recovery, and temporarily scaled-down environments. Forgotten gateway instances still matter if they retain network connectivity or credentials. An inactive interface does not guarantee an unreachable service.

A Custom Flow Turned Template Data Into Executable Behavior

The core failure was not that an AI model wrote a dangerous command. It was that a template boundary allowed crafted data to reach executable application behavior.

GitLab describes the issue as improper neutralization in a custom flow prompt template. A custom flow is a configurable, multi-step AI workflow within the Duo Agent Platform. It can combine prompts, components, routing decisions, and tools in a YAML definition.

The company’s custom flow documentation shows why these definitions are more than ordinary prompt text. Flows can be created, tested, published, enabled for projects, and triggered through GitLab activity. They operate as application configuration around an agent.

The vulnerable path involved a prompt template sandbox. A sandbox is a restricted execution environment intended to prevent a template from reaching unsafe attributes or functions. CVE-2026-90970 allowed a crafted flow configuration to escape those restrictions.

A public GitLab disclosure ticket attributes the issue to unsafe method access during Jinja2 template rendering. Jinja2 is a Python template engine that combines text with variables and expressions.

According to the ticket, conversation history contained a LangChain HumanMessage object. The template could call a public deserialization method on that object. An unsafe serialized payload then caused Python to execute operating-system commands during deserialization.

Deserialization converts stored or transmitted data back into a program object. It becomes dangerous when the selected format can invoke code while reconstructing that object. Python pickle data is a well-known example because loading untrusted pickle content is unsafe.

The reported proof of concept used a two-step flow. The first model interaction populated conversation history. A later template evaluation accessed that history object and triggered the dangerous deserialization path.

This sequence makes the issue different from a basic shell injection bug. The malicious value did not need to appear as a direct command parameter passed to a shell. Instead, several individually meaningful features formed the execution chain.

The flow accepted configurable prompt content. The template engine received application objects. One object exposed a deserialization method. The accepted serialization format could invoke code. Together, those conditions defeated the intended sandbox.

The model itself was not the trusted security boundary. It helped advance the flow between states, but the command executed through deterministic application behavior. Filtering model output alone would not address the vulnerable object and template relationship.

That distinction matters when organizations classify AI security failures. Prompt injection describes attempts to manipulate a model through instructions. CVE-2026-90970 is a software vulnerability in the system surrounding the model, even though a prompt template provides the entry point.

Traditional application-security practices therefore remain essential. Template inputs need strict validation, exposed objects need minimal interfaces, and dangerous deserialization formats should not process attacker-controlled data. Sandboxes also need testing against the exact objects made available inside them.

The reported commands ran with the privileges of the Duo Workflow Service process. That limits the immediate operating-system authority to the service account’s permissions. However, service-level execution is still serious because the gateway handles sensitive connections and sits within trusted infrastructure.

The practical impact depends on deployment architecture. A minimally privileged container with restricted networking presents fewer follow-on paths than a broadly connected service. Neither configuration removes the need to patch, because an attacker would still gain unintended code execution.

This mechanism explains the near-maximum severity. The attacker begins with an authenticated Duo-capable identity, but the resulting execution crosses into the gateway’s security context. That scope change is central to the 9.9 rating.

Self-Hosting Shifts Control and Security Responsibility Together

The incident exposes the tradeoff behind self-hosted AI infrastructure: keeping processing close also places the gateway inside the customer’s operational trust boundary.

GitLab describes the AI Gateway as a standalone service that connects GitLab Duo features with AI models. Customers can use GitLab’s managed gateway or operate their own gateway with GitLab Duo Self-Hosted.

Organizations often select self-hosting to control data movement, model access, network routes, and infrastructure policy. Those benefits can matter in regulated environments or deployments with strict internal controls. They also create another production service that customers must inventory and maintain.

The GitLab AI Gateway vulnerability turns that operational detail into the main security issue. GitLab could deploy a fix directly to gateways it manages. Self-hosted customers must schedule, execute, and verify their own upgrades.

This pattern is familiar from databases, identity services, and CI runners. The difference is that AI gateways join systems that were once more clearly separated. They sit between users, source repositories, agent workflows, model providers, and sometimes execution tools.

A compromised gateway therefore deserves more attention than an isolated chatbot interface. Depending on configuration, it can encounter prompt content, workflow state, service credentials, provider connections, or metadata about internal development activity.

That does not establish that CVE-2026-90970 exposed every connected secret. GitLab’s advisory does not report confirmed data theft or a complete post-exploitation path. The correct conclusion is that arbitrary command execution creates a credible route to further access.

Teams should evaluate the gateway as a privileged integration service. Its process identity, mounted files, environment variables, network routes, and attached service accounts determine the blast radius. Those controls become important when reconstructing exposure before the patch.

Containerization helps only when its boundaries are deliberately configured. A container can still reach network services, read mounted secrets, or send data outward. Its effective security depends on runtime permissions and surrounding policy.

GitLab’s installation guide recommends restricting outbound network access for the AI Gateway container. Egress control limits which destinations a compromised process can contact. It can reduce the usefulness of command execution, although it does not eliminate local impact.

Network segmentation provides another containment layer. A gateway needs access to defined GitLab and model endpoints, but it rarely needs unrestricted reach across an internal network. Narrow allowlists make unexpected lateral movement harder.

Credential design matters equally. Long-lived secrets stored directly in the environment create attractive post-exploitation targets. Short-lived credentials, isolated service accounts, and narrowly scoped permissions can reduce damage after a service compromise.

Security teams should also examine who can create or modify custom flows. GitLab’s documentation assigns flow-management actions to roles such as Maintainer or Owner in several workflows. Exact exposure still depends on product configuration and the affected implementation.

The advisory uses the broader phrase “Duo Agent Platform access” rather than naming one universally required project role. Administrators should not convert documentation examples into a definitive exploit prerequisite. They should review actual permissions and historical flow changes.

Self-hosting remains a valid architectural choice. The lesson is not that a managed service is always safer. The lesson is that data control, software control, and incident responsibility arrive together.

A managed gateway concentrates trust in the vendor’s operations. A self-hosted gateway concentrates trust in the customer’s patching, isolation, and monitoring. CVE-2026-90970 makes that exchange visible because the remediation boundary follows the hosting boundary exactly.

For enterprise buyers, security review should cover both product features and deployment ownership. Questions about where data travels should be paired with questions about who patches each component. An architecture diagram without operational ownership remains incomplete.

A Patch Closes the Flaw but Leaves Detection Questions

Updating stops the known vulnerable path, but the public advisory does not tell operators how to prove that earlier exploitation never occurred.

As of October 4, GitLab had not stated publicly that CVE-2026-90970 was being exploited in the wild. That absence is reassuring, but it is not evidence that every affected deployment remained untouched.

Public technical detail increases the importance of rapid patching. The disclosure ticket describes the vulnerable object relationship and reports successful command execution in a staging environment. Defenders and attackers can both study that material.

GitLab credited the HackerOne researcher known as invisiblemeerkat with responsible disclosure. Responsible reporting gave GitLab time to prepare fixes and contact affected customers. It did not remove the exposure window for deployments that remain unpatched after publication.

The authentication requirement should shape threat hunting. Security teams should start with Duo Agent Platform identities, flow-management events, and changes to custom flow definitions. They should correlate those records with gateway activity during the vulnerable period.

Unusual flow configurations deserve review, especially templates that access conversation-history objects or invoke methods. Unexpected flow creation, editing, publication, or execution can provide additional context. Normal-looking flow names should not override suspicious template behavior.

Gateway process activity also matters. Child processes, shell invocation, unexpected binaries, unusual file access, and outbound connections can indicate command execution. The useful data depends on container, host, and cloud telemetry already enabled.

Container restarts can erase local evidence. Centralized logs and runtime security telemetry are therefore more useful than inspecting only a currently running container. Teams should preserve available logs before replacing infrastructure if they suspect compromise.

Operators should review secrets accessible to the gateway process. Rotation decisions should follow evidence and exposure, not panic. If logs indicate command execution, assume that readable credentials might have been accessed.

The gateway’s connections to GitLab and model providers deserve separate attention. An attacker who gained command execution might attempt to reuse tokens, inspect configuration, or reach connected services. The public record does not confirm that such activity occurred.

The patch should not become the end of the investigation when suspicious evidence exists. Updating removes the known code path, but it does not revoke stolen credentials or undo changes made elsewhere. Incident-response procedures remain necessary.

There is also a historical reason for caution. Security reporting identified an earlier 2026 AI Gateway issue, CVE-2026-1868, with the same 9.9 rating and the same CWE-1336 weakness category. That flaw also involved crafted flow content and possible code execution.

Two high-severity template-engine issues do not prove that every custom flow is unsafe. They do justify closer review of how templates, application objects, and serialization meet inside agent systems.

The recurrence suggests that security testing needs to cover compositions, not only isolated components. A sandbox can behave correctly against strings while failing when rich framework objects enter its context. A safe method in one layer can become dangerous when templates can invoke it.

Agent platforms intensify this problem because they join many flexible mechanisms. Prompts, tools, histories, routing logic, model responses, and application APIs interact across repeated steps. State introduced in one step can become executable input later.

Organizations should add adversarial custom flows to predeployment testing. Tests should include method access, object traversal, serialization boundaries, and multi-step state changes. A one-pass prompt scan would not represent the reported exploit chain.

They should also treat templates as code-adjacent artifacts. Review, ownership, change history, and deployment controls should match their potential impact. Calling a file “configuration” does not reduce its ability to alter runtime behavior.

GitLab has not publicly detailed every condition required for exploitation. That limits confident exposure assessment. Teams should use the disclosed prerequisites as minimum conditions, not assume that unspecified conditions guarantee safety.

Three Signals Will Show Whether the Response Is Working

The next phase depends on patch adoption, evidence of real-world exploitation, and GitLab’s deeper treatment of custom-flow isolation.

The first signal is the percentage of self-hosted gateways running 19.2.4, 19.3.2, 19.4.1, or a later fixed version. Individual organizations should measure this across every environment and replica. Industry-wide numbers might remain unavailable because these deployments sit inside customer networks.

Fast adoption would narrow the reachable attack surface after public disclosure. Slow adoption would extend the risk, especially where AI services fall outside established vulnerability-management inventories. Gateway ownership should become visible in patch dashboards.

The second signal is any change in exploitation status. GitLab, CISA, incident-response firms, and affected customers might publish indicators or confirmed cases. A report of active exploitation would shift the priority from preventive patching toward broader incident response.

Defenders should distinguish a public proof of concept from observed attacks. Technical reproduction proves that the flaw works under documented conditions. It does not establish that attackers compromised production customers.

The third signal is a structural security change in the AI Gateway. A narrow patch can block the disclosed method call. A broader response might reduce which objects reach templates, prohibit dangerous serialization, or strengthen isolation around custom flows.

That design response matters because the reported chain emerged from feature composition. Preventing one payload is useful, but eliminating the unsafe capability boundary provides stronger protection against variants.

GitLab’s future release notes and code changes should clarify which layer received the fix. Administrators should watch for new guidance covering flow validation, audit events, detection queries, and supported upgrade paths for older gateway branches.

Enterprise customers can use the incident to test their own operating model now. The important question is not simply whether GitLab appears in the software inventory. It is whether the AI Gateway exists as a separately owned, patched, logged, and isolated service.

Developers who create flows should also reconsider the trust assigned to configuration. A custom flow can coordinate tools and application data across multiple steps. It should receive the same skepticism applied to automation scripts and CI definitions.

Security reviewers should map four boundaries: who can author flows, what objects templates can access, which tools flows can invoke, and what privileges the gateway process holds. Weakness across several boundaries can turn limited access into infrastructure control.

Knowledge workers using GitLab Duo do not need to abandon ordinary work because of this advisory. Most cannot determine gateway ownership themselves. They should follow organizational guidance and report unexpected flow behavior rather than attempting independent testing.

Administrators face a clearer action. Identify the hosting model, confirm the running gateway version, deploy the correct patch, and preserve evidence where suspicious activity appears. Review flow changes and gateway telemetry across the vulnerable period.

After patching, document the result in a durable operational record. Record the previous version, deployment time, affected environments, verification method, and any threat-hunting findings. That evidence supports future audits and incident reconstruction.

The GitLab AI Gateway vulnerability is ultimately a warning about where AI application logic becomes conventional executable software. The failure began in a prompt template, crossed through a framework object, and ended with operating-system commands.

That path deserves more attention than the “AI” label alone. Models can influence workflows, but ordinary software boundaries still decide whether untrusted input becomes code. Organizations need controls around both layers.

If your organization operates GitLab Duo, ask one concrete question today: who owns the gateway that processes those requests? If the answer is your team, verify its version against GitLab’s fixed releases. Then test whether your monitoring would reveal unexpected commands, altered flows, or outbound connections. The patch closes the disclosed GitLab AI Gateway vulnerability, but durable safety depends on inventory, isolation, and evidence that survive the next advisory.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page