top of page

Cloudflare Containers, Rebuilt to Scale Agent Sandboxes, Trades Static Deployment for Runtime Control

Oct 2
14 min read

Cloudflare Containers, rebuilt to scale agent sandboxes, now start more than six times faster, according to the company. The September 30 release also adds runtime image selection, runtime instance sizing, and filesystem snapshots in public beta.

The important change is not simply a shorter cold start. Cloudflare has moved control of each sandbox into a Durable Object, its stateful serverless component for coordinating requests and persistent application state. That decision challenges the static deployment model that shaped the first version of Cloudflare Containers.

Developers can now let an agent choose the environment required for each task. A coding job might need a larger instance and a full Linux toolchain. A smaller automation might use a lighter environment. When work pauses, the system can save the filesystem, stop the compute, and restore the workspace later.

That places Cloudflare more directly into a busy agent infrastructure market. E2B, Modal, Daytona, Vercel, and hyperscale cloud services already offer different approaches to isolated execution. Cloudflare’s argument is that a globally addressable controller, an isolated Linux workspace, and persistent state should operate as one programmable unit.

The architecture looks suited to coding agents, evaluations, and long-running workflows. However, the six-times startup claim comes from Cloudflare, snapshots remain in public beta, and several operational limits matter. The real test is whether teams gain reliable control without inheriting too much lifecycle complexity.

Cloudflare Containers, Rebuilt to Scale Agent Sandboxes, Changes the Control Plane

Cloudflare’s central change is moving sandbox configuration from deployment time to the moment an agent starts a task.

Under the earlier model, an application generally declared one container image and instance type in its deployment configuration. Changing either setting triggered an application-level rollout. That pattern works for predictable services, but agents generate workloads that vary from one request to another.

A code repair agent may need a repository, compiler, package manager, browser, and test suite. An evaluation worker may need a clean, repeatable environment with tightly controlled inputs. Another task may only require a short script with limited memory and no internet access.

Cloudflare’s new durable_object scheduling policy lets application code make those choices at runtime. The controlling Durable Object calls ctx.container.start() and supplies an image, snapshot, and instance configuration appropriate for that task.

The available predefined sizes include lite and four standard configurations. Developers can also pass custom CPU, memory, and disk values within platform limits. The runtime scheduling policy replaces one centrally selected configuration with per-sandbox decisions.

Image selection follows the same model. Developers declare named images through Wrangler, Cloudflare’s command-line deployment tool. The platform prepares immutable references, and the Durable Object selects one when starting a sandbox.

This arrangement lets one application support several agent roles without deploying a separate container application for every role. A coordinator can direct a lightweight research task toward one image, then assign a build job to another image with more resources.

Cloudflare also introduced cloudflare/debian-trixie, a managed Debian image containing Node.js. An agent can start with that foundation, install its tools through exec(), and preserve the resulting workspace as a snapshot.

According to Cloudflare’s sandbox announcement, the new scheduling path starts Containers more than six times faster. The company says it removed several coordination steps that previously sat between the Durable Object and the container runtime.

That claim needs careful interpretation. Cloudflare has not presented an independent benchmark comparing representative images, regions, or workload types. Container readiness also depends on image size, entrypoint behavior, capacity, and application-level health checks.

Cloudflare’s own architecture documentation says cold starts often fall within one to three seconds. It also notes that startup time varies with the image and its initialization work. A faster scheduler cannot eliminate delays created inside the image itself.

The platform exposes a running property before a process is necessarily ready to accept traffic. Developers must still check port readiness before sending the first request. For an agent, “container started” and “workspace ready for useful work” remain different measurements.

Even with those qualifications, runtime configuration changes the product’s operating model. Cloudflare Containers are no longer only deployed services that agents happen to use. They become resources that an agent controller can assemble, size, stop, and reconstruct around individual tasks.

Faster Agent Sandboxes Put Pressure on Static Deployment Models

Agent workloads reward infrastructure that can change shape between tasks, not merely infrastructure that runs one image efficiently.

Traditional container platforms assume developers know the application shape before deployment. Teams choose an image, resource allocation, networking policy, and scaling configuration. A scheduler then creates replicas that broadly share those properties.

Agent systems disturb that assumption. Their next action depends on user requests, model decisions, tool results, and the state left by earlier work. Two consecutive tasks within the same product may need different operating systems, dependencies, resource limits, and network permissions.

That variability creates pressure on providers built around static application configuration. Teams can still deploy several services and route work among them. However, each new workload type adds another deployment unit, rollout path, capacity decision, and source of configuration drift.

Cloudflare’s runtime model moves part of that decision into application code. The agent controller can choose a sandbox image and instance type after examining the task. It can also decide whether the sandbox receives internet access or starts from a stored snapshot.

The primary opponent is therefore not one named vendor. It is the static deployment model that treats every sandbox in an application as a copy of the same predefined service.

E2B, Modal, Daytona, and Vercel already address this market through their own abstractions. Some emphasize developer-friendly sandbox APIs. Others build around functions, virtual machines, workspaces, or broader cloud orchestration. Teams running Kubernetes or Firecracker directly gain more control, but they also own more infrastructure.

Cloudflare’s differentiator is the relationship between each container and its Durable Object. A Durable Object provides a stable identity, application code, storage, alarms, and coordination outside the Linux environment. The container supplies isolated compute for tools that need a conventional operating system.

The agent’s decision loop can remain active while the Linux workspace sleeps. It can communicate with users, retain authorization state, and call models without keeping the heavier sandbox running. When a compiler or development server becomes necessary, the controller wakes the container.

Cloudflare describes this separation as keeping the agent’s “brain” apart from its “hands.” The controller retains intent and state, while the sandbox performs commands that may fail, stop, or require replacement.

That separation has security benefits as well as operational benefits. The Durable Object can hold credentials outside the container and mediate outbound requests. An agent does not necessarily need direct access to every secret required for an authorized service call.

Cloudflare previously argued that dynamically generated agent code requires an isolated execution environment. Its earlier code sandbox model focused on lightweight Dynamic Workers for tasks that do not require a full Linux system.

Containers cover the heavier side of that split. They support package managers, native binaries, repositories, compilers, terminals, and development servers. Dynamic Workers can handle smaller code execution tasks with a narrower runtime.

This creates a tiered execution strategy. A controller can use a lightweight sandbox for a short API workflow and reserve a container for work requiring Linux. Runtime selection matters because keeping every task inside a full container wastes startup time and resources.

The pressure extends beyond sandbox vendors. Internal platform teams often maintain pools of warm development environments to hide provisioning delays. Faster cold starts and resumable filesystems weaken the case for keeping large idle pools available.

However, Cloudflare is not eliminating orchestration. It is relocating orchestration into the Durable Object and its application code. Teams must still design admission rules, concurrency controls, retry behavior, authorization, cleanup, and observability.

The winning model will not be whichever platform reports the smallest isolated startup number. It will be the model that minimizes total time from an agent’s decision to a verified task result.

That includes image preparation, repository access, dependency restoration, command execution, network latency, and teardown. It also includes the human delay caused by failed sessions or lost work.

For engineering teams comparing options, the relevant benchmark should reproduce their complete workflow. A synthetic empty-container test cannot represent a large repository, package installation, browser startup, or a test suite.

Those evaluations also produce design knowledge that teams need to retain. A searchable engineering knowledge base can preserve benchmark assumptions, security decisions, and migration findings alongside the implementation.

Durable Objects Turn Containers Into Task-Specific Compute

The mechanism behind the release is a stateful controller that treats its container as replaceable compute rather than permanent application state.

Every Cloudflare Container is associated with a Durable Object. Requests reach a Worker first, then route through that object before reaching the container. The Durable Object can address a specific workspace and preserve state connected to its identity.

With the new API, developers extend DurableObject directly and access the attached container through this.ctx.container. This removes the wrapper class that Cloudflare originally used to make Containers resemble conventional sandbox services.

The controller can start the container, execute commands, inspect it, monitor exits, send process signals, set an inactivity timeout, and destroy the instance. It can combine those controls with Durable Object storage, alarms, WebSockets, and remote procedure calls.

Consider a coding agent responding to a bug report. The Durable Object can store the session identifier, approved repository, user permissions, and current task phase. It can then select an image containing the appropriate language toolchain and start a suitable instance.

The container clones the repository, installs dependencies, runs the test suite, and edits files. Meanwhile, the Durable Object can send progress through a WebSocket and record checkpoints outside the container.

If the container process exits, the controller retains the session’s identity and metadata. It can inspect the failure, restart from a known state, or report the problem without losing the entire interaction.

Cloudflare runs each container inside a Firecracker microVM, a lightweight virtual machine with its own kernel and network. The customer image runs as a Linux container inside that virtual machine.

The platform’s container architecture states that other Cloudflare workloads do not share that kernel. This isolation matters because agent-generated commands should not run directly inside the application handling trusted user data.

Placement remains dynamic. Cloudflare chooses eligible capacity where the required image is available, with routing and startup speed influencing the location. The Durable Object and container are not guaranteed to run in the same place.

That caveat matters for latency-sensitive control loops. A globally addressable identity does not mean every operation runs beside the user, the model provider, or the sandbox. Teams should measure the full request path across their expected regions.

Capacity can also move between sessions. If a container stops and later restarts, Cloudflare may place the replacement elsewhere. Applications must not treat one machine’s local identity as permanent.

The Durable Object becomes the continuity layer. It stores the information required to find, rebuild, or restore the workspace. The Linux instance becomes an execution resource that can disappear when idle.

This architecture also supports branching workloads. A coordinator can start several independent attempts from the same prepared baseline. Each attempt can test a different model, system prompt, skill set, or repair strategy.

The controller can monitor those runs, compare results, and preserve the preferred output. Reinforcement learning systems can use similar patterns to create controlled environments, grade outcomes, and reset state between trials.

Runtime instance sizing strengthens that model. A controller can assign more resources to builds and reduce resources for lighter commands. Static application-wide sizing would force teams to provision for the largest common task or maintain separate deployments.

Yet programmability transfers responsibility to the application. The controller must prevent a model from selecting resources without limits. It should map agent requests onto approved policies rather than passing arbitrary model output into infrastructure APIs.

The same principle applies to images. Allowing runtime selection does not mean allowing an agent to execute any unreviewed image. Cloudflare requires declared, digest-pinned image references, which helps keep deployments reproducible.

Teams should still maintain an image allowlist, scan dependencies, restrict outbound access, and separate credentials from the guest environment. A sandbox reduces exposure, but it does not define the complete security policy.

Operationally, the Durable Object should remain the source of truth for lifecycle state. Cloudflare’s direct API provides control, but applications must decide when a task becomes recoverable, abandoned, completed, or safe to retry.

That is the real mechanism behind faster agent sandboxes. The scheduler improvement matters, but the lasting change is an explicit control plane that survives beyond any one container process.

Filesystem Snapshots Preserve Files, Not Running Sessions

Snapshots reduce repeated setup work, but they are immutable filesystem checkpoints rather than complete suspend-and-resume images.

Cloudflare’s native filesystem snapshots are available in public beta through the durable_object scheduling policy. A running container calls snapshotContainer() to capture its writable root filesystem at a point in time.

The returned handle contains an identifier, size, and optional name. Developers must save that handle, often in Durable Object storage, because the Worker API does not provide a command for listing snapshots.

A later container can start from the stored handle. This makes a coding workspace recoverable after the original compute stops. The repository, installed dependencies, build caches, configuration files, and edits can return with the restored filesystem.

That model addresses a common mismatch in agent infrastructure. Starting an empty sandbox may take seconds, while preparing a useful development environment can take minutes. Repeating dependency installation can dominate the startup measurement that users actually experience.

Snapshots also provide a stable baseline for evaluations. A team can prepare one repository and toolchain, save it, and start multiple experiments from the same checkpoint. Each sandbox receives an independent writable environment after restoration.

This reduces environment drift between attempts. If two model versions see different dependency states, test results become harder to compare. A shared immutable checkpoint helps isolate the variable under evaluation.

The feature also supports longer projects. An agent can save its workspace when a user leaves, stop the compute, and restore the files when the user returns. This separates filesystem continuity from continuous resource use.

However, the word “snapshot” can imply more than Cloudflare currently preserves. The snapshot documentation says the system captures the full container filesystem, but not memory, running processes, or separately mounted filesystems.

A restored container runs its entrypoint again. An in-memory build, live debugger, terminal process, or development server does not resume from the exact instruction where it stopped. The application must reconstruct those processes.

Snapshots are also tied to the image version used to create them. Developers cannot restore one into a different image. When a base image changes, the team must create a new compatible snapshot.

Each snapshot handle has an implicit 30-day time-to-live. Restoring it refreshes that period, but developers cannot configure a different retention interval today. That limitation makes snapshots unsuitable as an indefinite archive without an external preservation plan.

The snapshots are immutable. Changes made after restoration require another snapshot if the team wants to preserve them. Applications therefore need checkpoint policies that balance recovery value against storage, latency, and operational complexity.

Public beta status adds another reason for caution. Production teams should validate snapshot creation and restoration under interruption, concurrent access, image updates, and regional placement changes.

They should also test failure boundaries. If the container stops during a checkpoint, the application needs a clear record of which snapshot remains valid. If a task changes an external system, restoring its filesystem does not reverse that external action.

This distinction is especially important for autonomous agents. A rollback can restore local files while leaving a pull request, database update, email, or cloud resource unchanged. Retrying the task blindly could duplicate an irreversible action.

The controller therefore needs a task ledger outside the sandbox. It should record approved operations, external side effects, and completed checkpoints. The filesystem alone cannot represent the full truth of an agent’s work.

Security teams also need to examine what snapshots preserve. Repositories, generated code, logs, cached packages, and temporary credentials can all reach the writable filesystem. A snapshot may retain sensitive material longer than the running session.

Cloudflare offers outbound request interception and external credential handling, which can reduce secret exposure inside the guest. Developers should still verify that tools do not copy tokens into configuration files, shell histories, or package caches.

The company’s migration direction introduces another practical issue. New features, including the faster path and native snapshots, require direct ctx.container access. Cloudflare says it will maintain the older Container and legacy Sandbox classes through December 31, 2026.

Existing deployments will continue running after that date, according to the company, but those classes will stop receiving updates. Teams wanting the new capabilities must migrate their lifecycle logic to the direct Durable Object API.

That is a meaningful architecture change, not a flag that every team can enable safely. The direct API exposes more of the system’s identity and coordination model. It also asks developers to take explicit ownership of behavior previously hidden by a base class.

Cloudflare is turning Sandbox SDK 1.0 into a collection of utilities rather than a superclass. That should let teams combine convenient helpers with native lifecycle control. It also confirms that Cloudflare wants the Durable Object, not the SDK wrapper, to define the core abstraction.

The Next Tests Are Reliability, Adoption, and Competitive Response

The release becomes consequential only if real agent systems convert its new controls into lower task latency and reliable recovery.

The first signal to watch is production performance across complete workloads. Cloudflare’s reported six-times improvement concerns its startup path, but developers need measurements from task dispatch through application readiness.

Useful tests should cover small and large images, several regions, restored snapshots, fresh filesystems, and different instance sizes. They should distinguish scheduler latency from image initialization, dependency restoration, and service readiness.

If those measurements show consistently shorter end-to-end delays, Cloudflare’s argument strengthens. If gains disappear once repositories and development servers enter the workflow, startup speed becomes a narrower advantage.

The second signal is snapshot reliability after the public beta reaches demanding workloads. Teams should watch restoration failure rates, checkpoint duration, storage behavior, image upgrade procedures, and recovery after interrupted sessions.

Successful adoption by coding agents would be especially revealing. Cloudflare already cites Base44 and Kilo Code as users of isolated environments for agent work. Earlier Cloudflare material also identified Figma Make as a Containers customer for untrusted code execution.

These customer references show genuine use cases, but they do not provide comparative performance data. Independent engineering reports would carry more weight than launch testimonials.

Snapshot behavior under large repositories will also matter. Package directories and build outputs can expand quickly. Teams need to understand whether snapshots remain fast and manageable as workspaces grow across repeated checkpoints.

If Cloudflare expands snapshot retention controls, listing APIs, observability, or cross-version migration tools, that would suggest the feature is progressing toward broader production use. Persistent limitations would weaken the long-running workspace story.

The third signal is how competitors respond to Cloudflare’s combined controller and sandbox model. The agent sandbox market includes focused providers and broader compute platforms, each with different strengths.

E2B emphasizes purpose-built cloud environments for agents. Modal connects sandboxed compute to a larger serverless platform. Daytona focuses on development environments, while Vercel integrates sandbox capabilities with its application platform.

Hyperscale clouds give teams extensive control through virtual machines, containers, identity systems, and orchestration services. Their disadvantage is often the engineering work required to assemble those components into a low-latency agent product.

Cloudflare is betting that Durable Objects compress that assembly. Identity, state, communication, policy, and lifecycle can live beside the sandbox controller. Its global network supplies routing and prepared capacity.

A competitive response could take several forms. Other providers might add stronger stateful coordination, more flexible snapshots, finer runtime sizing, or tighter credential mediation. They could also publish benchmarks that challenge Cloudflare’s startup claims.

If competitors converge on an external stateful controller, Cloudflare’s architecture will look prescient even when customers select another platform. If developers favor simpler hosted sandbox APIs, the Durable Object model may feel too infrastructure-heavy.

Adoption will depend partly on how much control application teams want. A startup building a coding agent may value direct lifecycle programming. A team adding one code-execution feature may prefer an opinionated service with fewer decisions.

Cloudflare must serve both groups without hiding the capabilities that distinguish the platform. The new SDK direction tries to split that difference by offering utilities around a native API rather than another mandatory abstraction.

Security incidents will be another decisive measure. Agent sandboxes execute untrusted or unpredictable commands, often with network access and proximity to proprietary code. A fast environment that leaks credentials or permits cross-tenant access would fail its central purpose.

Cloudflare’s use of Firecracker microVMs provides kernel isolation between workloads. Still, secure operation also depends on image hygiene, outbound controls, authorization, secret handling, logging, and application policy.

Teams should treat the model as layered containment, not permission to trust agent-generated code. The Durable Object can become the policy enforcement point, but developers must implement and test that policy.

The release also raises a broader question about agent product architecture. Should the agent live outside the workspace and treat it as a replaceable tool, or should the agent run inside its computer?

Cloudflare supports both patterns. Running the agent in the Durable Object keeps communication and state available while compute sleeps. Running it inside the container offers a familiar Linux process model, while the object supervises it from outside.

The first pattern makes the brain-and-hands boundary explicit. The second may simplify existing agent runtimes that expect local files and processes. Actual adoption will show which model developers find easier to operate.

Cloudflare Containers, rebuilt to scale agent sandboxes, represents more than a cold-start improvement. It turns runtime choice, filesystem continuity, and lifecycle policy into application-level decisions controlled by durable state.

That shift puts pressure on static deployments and simplistic sandbox APIs. It also gives developers more ways to create subtle failures. Runtime flexibility requires strict policies, durable task records, and benchmarks that measure useful readiness instead of an empty process.

Teams evaluating the release should begin with one real workload. Measure fresh startup, restored startup, dependency readiness, failure recovery, and total task completion. Then test whether the controller survives lost containers without repeating external actions.

The next few months should reveal whether Cloudflare’s public beta snapshots remain dependable, whether customers publish independent latency results, and whether competitors adopt similar stateful controls. Those signals will determine whether this architecture becomes a common foundation for long-running agents or another specialized option in an increasingly crowded market.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page