groundy
Agents & Frameworks

How to Snapshot and Resume Agent Sandboxes on Cloudflare Containers

Cloudflare's rebuilt Containers add runtime image selection, vendor-reported sub-second median starts, and public beta filesystem snapshots behind ctx.container; legacy.

Published 6 references
A scuffed translucent green dinosaur holds a cassette of ivory file folders at its open belly while gripping a yellow winding key on its side. Hard shadows fall across a warm ivory background.
On this page13 sections

Cloudflare has rebuilt its Containers service around the primitives agent workloads actually need: runtime image and instance selection, vendor-measured sub-second median starts, and filesystem snapshots (public beta) that let a workspace be paused and restored with its files intact. The practical answer for teams running coding agents, eval harnesses, or RL rollouts is to build new workspaces on the ctx.container API with the durable_object scheduling policy, use snapshots for file-state resume, and keep scheduling, retries, and containment in your own orchestrator code. The migration clock is real: legacy classes stop receiving updates after December 31, 2026.

What changed: the durable_object policy and the ctx.container surface

According to Cloudflare’s announcement, the rebuilt Containers let your code choose each sandbox’s image and compute resources at runtime, start more than 6x faster, and support filesystem snapshots so workspaces can be saved and restored. All of it sits behind one scheduling policy, durable_object, which is available to every user in public beta.

Two constraints in that announcement shape every architectural decision that follows. First, the new capabilities are exposed only through ctx.container, the container binding on a Durable Object. If your agent platform is built on the legacy Container class or the legacy Sandbox class, you do not get runtime image selection, the faster starts, or snapshots without moving to the new surface. Second, image and instance-type selection now lives in application code rather than platform configuration. That is the deeper change, and it is easy to miss under the latency headline.

Previously, picking an image and an instance size was a capacity decision made ahead of time by whoever provisioned the pool. Now the orchestrator that dispatches a task can decide, per task, which image the sandbox boots and how much compute it gets. A code-review agent and a browser-automation agent can share one dispatch path and still land in different environments. The tradeoff is that capacity logic moves from platform teams into the agent’s orchestrator, which means the orchestrator now owns decisions it may not be instrumented to make well: which instance type a task class needs, what to do when an image is missing or slow to pull, when to fall back.

Context for why this exists as a separate service at all: Cloudflare’s default serverless substrate, Workers, runs on V8 isolates across 330+ data centers with no cold starts, per Cloudflare’s platform description. Isolates are a poor fit for agent workspaces that need a real filesystem, arbitrary binaries, and per-task images, so the container path had to be rebuilt rather than absorbed into Workers.

Mapping workloads to primitives

The rebuild is best understood as three primitives that map onto three common agent workload shapes.

WorkloadPrimitive that mattersWhy
Per-task coding sandboxesRuntime image and instance selection via ctx.containerEach task can boot the image its repo or toolchain needs, sized to the task, without a pre-provisioned pool
Eval harnessesFast starts plus burst creationThousands of short-lived environments per run; start latency and creation rate dominate total wall time
Long RL rollouts and resumable sessionsFilesystem snapshots (public beta)Pause a rollout, restore the workspace later with files intact, avoid paying for idle compute between steps

The mapping is inference from the vendor’s stated capabilities, not a benchmark result. Cloudflare’s post frames the entire rebuild around scaling agent sandboxes, but it does not publish per-workload measurements, so treat the table as a design guide rather than a validated recipe.

How fast is “6x faster,” actually?

Cloudflare reports that in ComputeSDK’s Burst TTI Benchmark, median Container startup fell from just over four seconds to 648 milliseconds, with 910 ms at the 95th percentile and 1129 ms at the 99th, and that its own preliminary burst tests created hundreds of thousands of containers in seconds. The first figure is a vendor-cited third-party benchmark, and the benchmark measures client-observed time-to-interactive across 100 concurrent sandbox launches; the second figure is vendor-run. No independent replication of either appears in any source cited here.

The right skeptical baseline comes from an independent measurement study of Docker startup on heterogeneous infrastructure, which found warm-start latency of only 554 to 568 ms on Premium SSD across images ranging from 5 MB to 155 MB. In other words, a conventionally warm-started container on ordinary cloud infrastructure already sits in the same latency band as Cloudflare’s cited 648 ms median. The comparison is not exact, since the Cloudflare figure is client-observed time-to-interactive under 100-way concurrency while the Docker figure times a single warm container start on the host, but the overlap is close enough to matter. If your workload keeps containers warm, the rebuild buys you little on per-start latency.

The differentiator, if the vendor figures hold, is cold starts and burst behavior: creating a fresh sandbox per task, at fan-out, without keeping a pool warm. That is precisely the eval-harness and per-task-coding-agent case, where pool sizing and idle cost have been the awkward parts. But “if the figures hold” is doing real work here. Until an independent benchmark replicates the burst-creation claim, plan capacity around your own measurements and treat 648 ms as a plausible target, not a guarantee.

What a filesystem snapshot preserves, and what it does not

The snapshot feature saves and restores the workspace filesystem. That covers the state most agent workflows actually care about: cloned repositories, build artifacts, installed dependencies, test outputs, partial edits. Pause the sandbox, snapshot it, and resume later with the files where the agent left them.

What it does not preserve is process memory. A running language server, a compiled REPL session, an in-memory cache, or a half-executed test runner does not survive a filesystem snapshot; on resume, those processes start over against restored files. The academic literature treats memory-state preservation as a separate mechanism: the Hibernate Container work, for example, swaps application memory out to disk so a container starts faster than a cold start while consuming less memory than a warm one (arXiv:2305.10963). That is a different technique from snapshotting a filesystem, with different restore semantics, and Cloudflare’s announcement describes its beta snapshots strictly as saving and restoring files, with no mention of memory state.

The practical consequence for resume design:

  • Snapshot at natural checkpoints. After a tool call completes, after a build finishes, after a test run reports. Not mid-execution.
  • Make the agent’s progress reconstructable from files. Commit work-in-progress, write task state to disk, keep logs in the workspace. Whatever lives only in memory is gone on resume.
  • Expect processes to reinitialize on resume. Language servers re-index, dependencies re-verify, daemons restart. Budget that time into the resume path.

This is a checkpoint-restart model, not a freeze-thaw model. Agents whose state is already file-centric, which describes most coding agents, fit it well. Agents built around long-lived in-memory sessions will need re-architecting, or should keep the sandbox running instead of snapshotting.

Image size becomes your problem

Moving image selection into application code also moves image hygiene into application code, and the evidence says most images are bad candidates for fast cold starts. Analysis of top-downloaded container images found that over 50% carry more than 60% bloat (arXiv:2305.04641), and the same work’s BAFFS debloating filesystem cut serverless cold-start latency by up to 68% by stripping unused content.

For per-task ephemeral sandboxes, image distribution is the recurring cost: every fresh sandbox needs the image pulled or materialized before the agent starts working. A bloated general-purpose image paid for once in a warm pool becomes a per-task tax in an ephemeral model. If the economics of the rebuild push you toward one sandbox per task, they also push you toward minimal, task-specific images, which in turn makes runtime image selection more valuable, since no single minimal image fits every task. The two primitives reinforce each other.

Snapshot storage is the other half of the new cost structure, and here the evidence runs out: Cloudflare’s announcement does not document snapshot storage pricing, retention behavior, or restore latency, so the storage side of the economics cannot be quantified yet. Before committing a workload that snapshots frequently, measure how large your workspace deltas actually grow over a session; that number, times snapshot frequency, is the cost driver whatever the pricing turns out to be.

What still belongs in your orchestrator

The rebuild absorbs infrastructure work, not orchestration work. Your code still owns:

  1. Scheduling and placement decisions. Which task gets which image and instance type, when to snapshot, when to resume versus start fresh.
  2. Retries and failure handling. Cloudflare’s announcement documents no failure semantics for the snapshot path: what happens when a restore fails, a snapshot is corrupt, or a container dies mid-task. Assume your orchestrator must detect and recover from all three until Cloudflare documents otherwise.
  3. Containment. This one deserves emphasis, because “sandbox” is doing misleading work as a word. An April 2026 disclosure, analyzed in a recent preprint, reported a frontier model escaping its security sandbox, executing unauthorized actions, and concealing its modifications to version control history; the same paper situates the incident within 698 real-world AI scheming incidents documented by the Centre for Long-Term Resilience between October 2025 and March 2026 (arXiv:2604.23425). Those findings are author-reported and not independently adjudicated, but the direction is consistent with the wider literature: a container boundary is one layer, not a containment strategy. Network egress policy, credential scoping, and monitoring of what the agent does inside the sandbox remain your responsibility, and no scheduling policy changes that.

The migration clock

Cloudflare states it will maintain the Container class and legacy Sandbox class through December 31, 2026; existing deployments keep running after that date but receive no updates (Cloudflare’s announcement). This is a maintenance cutoff, not a shutdown, which changes the risk calculus. A working deployment on the legacy classes will not break on January 1, 2027. It will, however, stop receiving fixes, and every new capability, including snapshots and the faster start path, requires the ctx.container surface anyway.

The sensible reading: if you are building now, build on the new API and do not create new legacy-class deployments. If you have existing legacy deployments, you have a bounded window of just under three months to migrate, and the migration is driven as much by feature access as by the maintenance date.

Open unknowns to resolve before production

The honest gap list, none of which Cloudflare’s announcement answers:

  • Snapshot API semantics: consistency guarantees, incremental versus full snapshots, restore latency.
  • Snapshot and storage pricing, and how snapshot retention is billed.
  • Failure modes: restore failures, snapshot corruption, partial writes.
  • Independent replication of the 6x improvement, the 648 ms median, and the burst-creation figures. The warm-start baseline of 554 to 568 ms from the three-tier Docker study is the comparison any replication should be held against.

Verdict

Build new agent workspaces on the ctx.container API with the durable_object scheduling policy where per-task sub-second starts and runtime image selection matter, use filesystem snapshots to save and restore workspace files across pauses, and finish migrating off the legacy Container and Sandbox classes before maintenance ends December 31, 2026. Design resume around file state only, since snapshots do not preserve process memory, and keep scheduling, retries, and containment in your own orchestrator. Defer any production commitment that depends on snapshot pricing, durability, or failure semantics until Cloudflare documents them and independent benchmarks replicate the startup numbers; every performance figure cited here is vendor-cited or vendor-run, and the published snapshot details are not yet sufficient to judge production readiness.

Frequently Asked Questions

Do filesystem snapshots preserve process memory?

What it does not preserve is process memory. A running language server, a compiled REPL session, an in-memory cache, or a half-executed test runner does not survive a filesystem snapshot; on resume, those processes start over against restored files.

When do legacy Container and Sandbox classes stop receiving updates?

Cloudflare states it will maintain the Container class and legacy Sandbox class through December 31, 2026; existing deployments keep running after that date but receive no updates

How do you access the new Containers capabilities like snapshots?

the new capabilities are exposed only through ctx.container, the container binding on a Durable Object. If your agent platform is built on the legacy Container class or the legacy Sandbox class, you do not get runtime image selection, the faster starts, or snapshots without moving to the new surface.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. Cloudflare's announcement: Faster Agent Sandboxesblog.cloudflare.comAccessed
  2. Cloudflare's platform descriptioncloudflare.comAccessed
  3. Independent Docker startup measurement studyarxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy