# Cloudflare Containers agent sandboxes: startup and setup

> Cloudflare Containers now start agent sandboxes in a median 648 ms. Set up the durable_object policy, runtime images and snapshots, and know the limits.

By Boris Tane, Founder of Polylane · Published October 1, 2026 · 15 min read
Canonical: https://polylane.com/learn/deployment-safety/cloudflare-containers-agent-sandboxes-startup-time/

On 30 September 2026 Cloudflare rebuilt Containers for agent sandboxes. In ComputeSDK's independent benchmark, median startup fell from 4.049 seconds to 648 milliseconds on the new durable_object scheduling policy. To use it, create a new Container application with scheduling_policy set to durable_object and start each sandbox from Durable Object code with ctx.container.start(), picking the image and instance size per task. The policy and filesystem snapshots are both in public beta, and you can't switch an existing application over, so plan a cutover.

Cloudflare rebuilt Containers for agent sandboxes and announced it on 30 September 2026. The change that cuts startup time is a new `durable_object` scheduling policy. In ComputeSDK's independent Burst TTI Benchmark, median time-to-interactive fell from 4.049 seconds to 648 milliseconds, and the 99th percentile fell from 6.717 seconds to 1129 milliseconds. The same policy lets your Durable Object pick each sandbox's image and instance type at start time. It also enables filesystem snapshots, which are in public beta.

To get the faster path, create a new Container application with `scheduling_policy` set to `durable_object`. Start sandboxes from Durable Object code with `ctx.container.start()`, and boot either the prepared `cloudflare/debian-trixie` image or a saved snapshot. The policy itself is in public beta and can't be added to an existing application, so plan a cutover. This page covers what changed, the setup step by step, and the limits Cloudflare's docs state.

## What is the durable_object scheduling policy?

A scheduling policy decides where a Container's image and instance size are set, and how image updates reach running instances. You choose it when you create the Container application, and you can't change it afterwards.

On Cloudflare, every Container has always been paired with its own Durable Object. That Durable Object is a stateful controller with a stable identity, and it starts, sleeps and stops the Container. Until this release, though, the image and the compute size were fixed at deploy time. Each pair of image and instance type was its own Containers application with its own Durable Object namespace, set up with `wrangler deploy`. If you needed a small Node.js sandbox and a large Python build sandbox, that meant two applications, two namespaces and routing logic in your Worker.

The two policies now work like this:

- **`default`**: Wrangler configuration holds one `image`, one `instance_type` and settings such as `max_instances`. Cloudflare applies changes across the whole application with rollouts. If you leave out `scheduling_policy`, you get this one.
- **`durable_object`** (public beta): Wrangler declares a map of named images the Durable Object may use. Each time your code boots a sandbox, it passes an image or a snapshot, plus an optional instance size, to `ctx.container.start()`.

One Wrangler configuration can hold applications with both policies. A Worker can keep service Containers on the `default` policy next to per-task sandboxes on `durable_object`.

The new policy suits workloads where the task decides the environment. Cloudflare lists four: coding agents that need repositories, package managers, compilers and test runners; evals that must start from a known state; reinforcement learning systems that create, grade and reset many environments; and long-running tasks that need to keep the files an agent produces. The announcement names Base44 and Kilo Code as users. It also lists integrations with Cursor Cloud Agents, Devin Outposts, the OpenAI Agents API and Claude Managed Agents.

## Where the faster startup comes from

**648 ms median time-to-interactive on the durable_object scheduling policy**

ComputeSDK Burst TTI Benchmark, 100 sandboxes launched at once, as published by Cloudflare. The previous scheduling path took 4.049 seconds.

ComputeSDK's Burst TTI Benchmark launches 100 sandboxes at the same time and measures time-to-interactive from the client. Cloudflare published these results for the old and new paths: the median went from 4.049 seconds to 648 milliseconds (6.2x faster), the 95th percentile from 5.839 seconds to 910 milliseconds (6.4x), and the 99th percentile from 6.717 seconds to 1129 milliseconds (5.9x).

**ComputeSDK Burst TTI results as published by Cloudflare**

| Startup measurement | Previous scheduling path | New scheduling policy | Improvement |
| --- | --- | --- | --- |
| Median | 4.049 seconds | 648 milliseconds | 6.2x faster |
| 95th percentile | 5.839 seconds | 910 milliseconds | 6.4x faster |
| 99th percentile | 6.717 seconds | 1129 milliseconds | 5.9x faster |

In its own preliminary burst test, Cloudflare started 100,000 Containers from a single account in 5.387 seconds across six locations. It calls that test preliminary, so measure your own workload before you depend on it.

The time comes out of three places:

1. **Placement starts at the Durable Object.** The old path went through a global control plane, which resolved the application configuration, found capacity and coordinated placement. Now the infrastructure looks for capacity on the same machine as the Durable Object first, then widens the search within the same location.
2. **Hosts that already hold the image get first pick.** The scheduler favours hosts that have the Container's image or snapshot in local storage, so the start can skip the download.
3. **The runtime restores a prepared virtual machine.** It restores an unassigned, prepared VM, reuses networking and filesystem setup, batches repeated operations, and no longer waits on services the first command doesn't need.

The benchmark measures the platform. Your time to a useful sandbox also includes image preparation and workspace setup such as cloning and installing dependencies. That is why the managed image and snapshots matter as much as the scheduler.

## What changes for a team on the default policy

If you run Containers today, these are the differences that will touch your code and your operations:

- **Configuration moves into code.** The image and instance size are arguments to `ctx.container.start()`. Adding a new environment means adding an image to the map and an `if` branch.
- **Rollouts move into code.** A running Container keeps the image it started with until your code stops it. You no longer configure grace periods or percentage splits, or push configuration through the API.
- **`max_instances` goes away.** Running instances count towards your account limits. If you want a cap for each application, enforce it in your own code.
- **Registries narrow.** Named images come from a Dockerfile or from a digest-pinned reference in the Cloudflare managed registry. Direct references to Docker Hub, Amazon ECR and Google Artifact Registry are not supported, and credentials set with `wrangler containers registries configure` are not used.
- **Snapshots only work here.** Applications on the `default` policy can't create or restore them.
- **The `Container` class doesn't support this policy.** Use the `ctx.container` API directly, and rebuild any helpers you relied on, such as port readiness checks, request proxying and sleep timeouts. The announcement says Cloudflare is carrying this model into Sandbox SDK 1.0.
- **Some settings drop out.** Placement constraints, rollout settings, `wrangler_ssh` and `trusted_user_ca_keys` are not supported. The `ssh` and `authorized_keys` fields still work.

## Starting your first agent sandbox, step by step

### 1. Declare the application and its images

Add a Container entry with the policy, a SQLite-backed Durable Object class and a binding. This example declares two images built from local Dockerfiles:

```jsonc
// wrangler.jsonc
{
  "name": "agent-computer",
  "main": "src/index.ts",
  "compatibility_date": "2026-09-29",
  "containers": [
    {
      "class_name": "AgentComputer",
      "scheduling_policy": "durable_object",
      "images": {
        "node": { "dockerfile": "./images/node/Dockerfile" },
        "python": { "dockerfile": "./images/python/Dockerfile" }
      }
    }
  ],
  "durable_objects": {
    "bindings": [{ "name": "AGENT_COMPUTER", "class_name": "AgentComputer" }]
  },
  "exports": {
    "AgentComputer": { "type": "durable-object", "storage": "sqlite" }
  }
}
```

Each key under `images` becomes a property on `ctx.container.images`, holding a digest-pinned reference. Each entry takes exactly one source: `dockerfile`, with optional `build_context` and `build_vars`, or `image`, which must point at `registry.cloudflare.com` with a `sha256` digest. A configuration can hold up to 100 named images, and each name must be between 1 and 128 characters long.

If you don't need a custom image yet, skip the map. Pass `cloudflare/debian-trixie` to `start()`. It is a Cloudflare-managed identifier for Node.js 24.20.0 on Debian Trixie slim, and Cloudflare distributes and prepares it on eligible hosts before requests arrive, so starts don't wait to download or unpack the base image.

### 2. Start a sandbox sized for the task

```ts
// src/index.ts
import { DurableObject } from "cloudflare:workers";

type Task = { runtime: "node" | "python" | "bare"; heavy: boolean };

export class AgentComputer extends DurableObject<Env> {
  startFor(task: Task) {
    if (this.ctx.container.running) return;

    const images = this.ctx.container.images;
    const image =
      task.runtime === "node" ? images.node :
      task.runtime === "python" ? images.python :
      "cloudflare/debian-trixie";

    this.ctx.container.start({
      image,
      instance: task.heavy ? "standard-3" : "standard-1",
      enableInternet: true,
    });
  }
}
```

`instance` accepts `lite`, `standard-1`, `standard-2`, `standard-3` and `standard-4`. If you leave it out, you get `lite`, which has 256 MiB of memory. For a custom size, pass an object such as `{ vcpu: 1, memoryMib: 4096, diskMb: 8000 }`. Note the camel case at runtime, where Wrangler's `default` policy uses `memory_mib` and `disk_mb`. Whenever you pass options, `enableInternet` is required.

### 3. Wait until the sandbox can take work

`start()` returns before the Container is ready to accept requests, and `running` being true doesn't confirm readiness either. Add your own readiness check. If your image runs a server, poll it over `getTcpPort()`:

```ts
async waitUntilReady(port = 8080, attempts = 50) {
  const target = this.ctx.container.getTcpPort(port);
  for (let i = 0; i < attempts; i++) {
    try {
      const res = await target.fetch("http://container/health");
      if (res.ok) return;
    } catch {
      // not listening yet
    }
    await new Promise((r) => setTimeout(r, 100));
  }
  throw new Error("sandbox did not become ready");
}
```

The health path and port belong to your image. With a bare image such as `cloudflare/debian-trixie`, you can retry a trivial `exec()` until it succeeds.

### 4. Run the agent's commands with exec()

`exec()` starts a process inside a running Container. It runs the executable directly, so pipes, redirects and `&&` need an explicit shell:

```ts
const proc = await this.ctx.container.exec(
  ["sh", "-c", "node --version && ls /workspace"],
  { cwd: "/", stderr: "combined" },
);
const out = await proc.output();
console.log(out.exitCode, new TextDecoder().decode(out.stdout));
```

`exec()` has no built-in timeout. To stop a runaway command, call `proc.kill()`, which sends `SIGTERM` by default, then watch `proc.exitCode`. A process can ignore the signal, so add your own deadline and escalate with `destroy()` if you have to. If a command produces large output, read both streams as they arrive. `output()` buffers everything and can only be called once.

### 5. Save the workspace and restore it later

A snapshot captures the Container's whole filesystem, so the repository, installed dependencies, build caches and the agent's edits come back without rebuilding. The handle is plain data, and you can store it in the Durable Object's SQLite storage:

```ts
async saveWorkspace() {
  const snap = await this.ctx.container.snapshotContainer({ name: "end-of-session" });
  await this.ctx.storage.put("workspace", snap);
}

async resumeWorkspace() {
  const snap = await this.ctx.storage.get<ContainerSnapshot>("workspace");
  if (!snap) return false;
  this.ctx.container.start({ containerSnapshot: snap, enableInternet: true });
  return true;
}
```

`image` and `containerSnapshot` are mutually exclusive, because a snapshot already identifies the filesystem. Snapshots are immutable and reusable. Several Containers can start from one snapshot and each make their own changes. That is the pattern for evals that need the same repository, dependencies and inputs across different prompts, models or agent versions. After a restored sandbox changes files, take a new snapshot to keep them.

### 6. Let idle sandboxes stop

By default, Cloudflare stops a Container shortly after its Durable Object becomes inactive. To keep a sandbox warm between an agent's requests, set an inactivity timeout of up to 6 hours. If the Durable Object handles another request before the timeout ends, the same Container is still running. Otherwise, Cloudflare stops it. Each Durable Object sets its own timeout and starts without one after a restart, such as a deploy, so set it after `start()` and again in the constructor when the Container is already running. Watch for exits too:

```ts
await this.ctx.container.setInactivityTimeout(10 * 60 * 1000);
this.ctx.waitUntil(
  this.ctx.container.monitor()
    .then(() => console.log("sandbox exited"))
    .catch((err) => console.error("sandbox errored", err)),
);
```

If the workspace matters, snapshot it before the timeout can fire. Durable Object alarms let you schedule that work without keeping the Container awake.

### 7. Deploy

```bash
npx wrangler deploy
```

Wrangler builds each Dockerfile image, uploads it and records the image map with the Worker version. Updating the map later doesn't restart or replace running Containers.

## Rolling out a new toolchain from Durable Object code

With this policy, a rollout is a choice your code makes at each start. Declare the candidate as a second named image, for example `nodeNext`, then pick between the two:

```ts
function bucket(id: string): number {
  let h = 0;
  for (const c of id) h = (h * 31 + c.charCodeAt(0)) >>> 0;
  return h % 100;
}

chooseImage() {
  const images = this.ctx.container.images;
  // Canary: 5% of sandboxes, chosen by Durable Object ID
  return bucket(this.ctx.id.toString()) < 5 ? images.nodeNext : images.node;
}
```

The announcement lists four patterns, and each is a few lines of code:

- **Canary** a toolchain on 5% of new sandboxes by hashing the Durable Object ID, as above.
- **Pin** active projects to their current image, so an agent never has its environment swapped mid-task.
- **Migrate** a workspace at a natural checkpoint, such as the next session or after a snapshot.
- **Roll back** by changing which image future starts choose. There is no configuration push and no drain.

To upgrade a running Container straight away, compare its image with the configured one. The docs warn that `inspect()` reports an empty image while a Container is starting and for one restored from a snapshot, so guard against that case:

```ts
const image = this.ctx.container.images.node;
const info = await this.ctx.container.inspect();
if (info && info.image !== "" && info.image !== image) {
  await this.saveWorkspace(); // optional; the snapshot stays tied to the old image
  await this.ctx.container.destroy();
}
if (!this.ctx.container.running) {
  this.ctx.container.start({ image, instance: "standard-2", enableInternet: true });
}
```

During a Workers gradual deployment, each Durable Object sees the image map of the Worker version that runs it. Different sandboxes can start different images until the deployment completes, so your canary logic should allow for that.

## Moving an existing Container application to the new policy

You can't flip `scheduling_policy` on an existing application. Create a replacement and cut traffic over:

1. **Add a new Durable Object class, binding and Container entry.** Give the entry a different `name`, set `"scheduling_policy": "durable_object"`, and point `class_name` at the new class. Leave the old entry, class and binding unchanged so you can route back. Declare the class the same way the Worker already does, through `exports` or the legacy `migrations` array.
2. **Port the class off `Container`.** Rewrite it against `ctx.container` and move `env`, `entrypoint`, the image and the instance size into the `start()` call. The `Container` class allows outbound Internet access unless told otherwise, so set `enableInternet: true` if you need the same behaviour.
3. **Map instance types.** `dev` becomes `lite`, and `standard` becomes `standard-1`. For `basic`, choose `lite` or `standard-1`. A custom instance needs at least 1 vCPU, so it can't reproduce `basic`.
4. **Move external images.** Push anything on Docker Hub, ECR or Artifact Registry to the Cloudflare managed registry, then reference it by digest.
5. **Deploy both and test behind a test-only route** that uses the new binding. Check the image, size, environment, network access and readiness before you move production traffic.

This process doesn't transfer running instances or Durable Object storage. Because the `default` policy can't take snapshots, you can't carry a filesystem across with one either. If your old Durable Objects hold data, design a transfer step for it first.

## Limits Cloudflare states for sandboxes

These are the predefined instance types, with whether `ctx.container.start()` accepts each name:

**Container instance types from Cloudflare's limits page**

| Instance type | vCPU | Memory | Disk | Accepted by start() |
| --- | --- | --- | --- | --- |
| lite | 1/16 | 256 MiB | 2 GB | Yes (the default) |
| basic | 1/4 | 1 GiB | 4 GB | No |
| standard-1 | 1/2 | 4 GiB | 8 GB | Yes |
| standard-2 | 1 | 6 GiB | 12 GB | Yes |
| standard-3 | 2 | 8 GiB | 16 GB | Yes |
| standard-4 | 4 | 12 GiB | 20 GB | Yes |

- **Custom instances:** 1 to 4 vCPU, at most 12 GiB of memory and 20 GB of disk, and at least 3 GiB of memory per vCPU.
- **Account limits:** 6 TiB of concurrent memory, 1,500 concurrent vCPU, 30 TB of concurrent disk and 50 GB of total image storage. An image can be as large as the instance's disk. With no `max_instances`, these are your ceiling.
- **Snapshots:** up to 20 GB each, kept for 30 days from creation or the most recent restore. Each restore resets that 30-day clock, and you can't set a custom retention yet.
- **Named images:** up to 100 for each configuration.

If you need larger sizes or higher account limits, Cloudflare's limits page says to contact your account team or file a support ticket.

## What goes wrong when adopting faster sandboxes

- **Editing the policy on the live entry.** `wrangler deploy` ships the new Worker version before it configures the Container application. The deploy then fails with the new code live and no `durable_object` application behind it. Fix: add a new entry and class, as in the migration steps.
- **Sending work the moment `start()` returns.** The call returns before the sandbox is ready. Fix: always run a readiness check before `exec()` or proxying traffic.
- **Shell syntax in `exec()`.** `["npm ci && npm test"]` fails because nothing interprets it. Fix: wrap it in `["sh", "-c", "..."]`, or use `bash -lc` if Bash is in the image.
- **Old instance names.** The runtime rejects `basic`, `dev` and `standard`. Leaving `instance` out gives you `lite`, which is too small for most builds. Fix: always pass a size.
- **Restoring a snapshot onto a new image.** A snapshot is tied to the image version it came from. Fix: after an image update, take fresh snapshots from Containers running the new image, and keep old sandboxes on their old image until they reach a checkpoint.
- **Expecting processes back after a restore.** Snapshots hold the filesystem only, with no memory or running processes. Fix: restart dev servers and watchers from your resume code.
- **Losing work with `destroy()`.** It stops the Container immediately and drops every filesystem change since startup. Fix: snapshot first when the files matter.
- **Snapshots expiring quietly.** A workspace left alone for more than 30 days is gone. Fix: store the handle with a timestamp, and fall back to a fresh image plus setup when the restore fails.
- **Docker Hub references in the image map.** These aren't supported on this policy. Fix: push the image to `registry.cloudflare.com` and pin it by digest, or build it from a Dockerfile whose `FROM` line pulls the external base.
- **Runaway fleets.** With no `max_instances`, a bug that starts sandboxes in a loop runs until it hits account limits. Fix: count active sandboxes in your own code, for example in a coordinating Durable Object, and refuse new starts above your cap.

## What to record about every sandbox start

I built and led the Workers observability team at Cloudflare. The habit I'd bring to this release is to log the decision your code makes at every start, because the platform no longer makes it for you. Once the image and size are chosen per request, "which sandboxes ran the new toolchain" lives only in your logs.

For each start, record the Durable Object ID, the image reference (or snapshot name), the instance size, `enableInternet`, and the time from `start()` to your first successful readiness check. Add the exit or error from `monitor()`. With those fields, you can compare your own startup times against the published benchmark, see whether a canary image starts slower, and tell a failed restore from a slow one.

Running on Cloudflare? See [how Polylane monitors Cloudflare in production](https://polylane.com/for/cloudflare/).

## Common questions

**When did Cloudflare announce the faster Containers for agent sandboxes?**

Cloudflare announced it on 30 September 2026 in the post Cloudflare Containers, rebuilt to scale agent sandboxes, during Birthday Week. The durable_object scheduling policy and filesystem snapshots are both in public beta. The scheduling policy docs were updated the same day.

**How much faster do Cloudflare Containers start now?**

In ComputeSDK's Burst TTI Benchmark, which launches 100 sandboxes at once, median time-to-interactive fell from 4.049 seconds to 648 milliseconds, a 6.2x improvement. The 95th percentile went from 5.839 seconds to 910 milliseconds, and the 99th from 6.717 seconds to 1129 milliseconds. These gains apply only to applications on the durable_object scheduling policy.

**Can I switch my existing Container application to the durable_object policy?**

No. The scheduling policy is fixed when you create the application. You add a new Container entry with a new Durable Object class and namespace, port your code to ctx.container, deploy both, and cut traffic over. Running instances and Durable Object storage don't transfer.

**Do I need a Dockerfile to start an agent sandbox?**

No. Pass cloudflare/debian-trixie as the image to ctx.container.start(). It is a Cloudflare-managed image with Node.js 24.20.0 on Debian Trixie slim, prepared on eligible hosts before requests arrive. Your agent can then use exec() to set up the environment it needs.

**What does a Cloudflare Containers snapshot save?**

It saves the complete filesystem of a running Container, up to 20 GB. Memory and running processes are not saved. A snapshot is tied to the image version it came from and is kept for 30 days from creation or the latest restore. It works only on the durable_object policy.

**Can I use Docker Hub or Amazon ECR images with the durable_object policy?**

Not directly. Named images must come from a Dockerfile or a digest-pinned reference in the Cloudflare managed registry at registry.cloudflare.com. Push external images there first. A Dockerfile can still use an external image in its FROM line.

**Which instance sizes can my code choose at runtime?**

ctx.container.start() accepts lite, standard-1, standard-2, standard-3 and standard-4, and it uses lite if you leave instance out. You can also pass a custom object with 1 to 4 vCPU, up to 12 GiB of memory and 20 GB of disk, and at least 3 GiB of memory per vCPU. The basic, dev and standard names are rejected at runtime.

**Does the Container class work with the new scheduling policy?**

No. The Container class doesn't support the durable_object policy, so your Durable Object must use the ctx.container API directly. You'll need to rebuild helpers such as readiness checks, request proxying and sleep timeouts. Cloudflare says it is carrying the same model into Sandbox SDK 1.0.

## Sources

- [Cloudflare Containers, rebuilt to scale agent sandboxes (Cloudflare Blog)](https://blog.cloudflare.com/faster-agent-sandboxes/)
- [Scheduling Policies (Cloudflare Containers docs)](https://developers.cloudflare.com/containers/configuration/scheduling-policy/)
- [Use snapshots (Cloudflare Containers docs)](https://developers.cloudflare.com/containers/guides/snapshots/)
- [Limits and Instance Types (Cloudflare Containers docs)](https://developers.cloudflare.com/containers/platform/limits/)
- [Migrate to the Durable Object scheduling policy (Cloudflare Containers docs)](https://developers.cloudflare.com/containers/guides/migrate-to-durable-object-scheduling-policy/)
- [Durable Object Container API (Cloudflare Containers docs)](https://developers.cloudflare.com/containers/api/durable-object-container/)
- [Image Management (Cloudflare Containers docs)](https://developers.cloudflare.com/containers/guides/image-management/)

## About the author

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

## Related

- [How to Debug Cloudflare Workers Errors: Logs, Traces and Error Codes](https://polylane.com/learn/troubleshooting/how-to-debug-cloudflare-workers-errors/): Debug Cloudflare Workers errors step by step: read 1101 and 1102 codes, enable Workers Logs and source maps, use wrangler tail, DevTools and local traces.
- [How to keep AI coding agents from breaking production](https://polylane.com/learn/ai-in-production/how-to-keep-ai-coding-agents-from-breaking-production/): Stop AI coding agents breaking production: scoped credentials, branch protection they can't bypass, required checks that block, and production-aware review.
- [Turn your app into a context graph](https://polylane.com/blog/turn-your-app-into-a-context-graph/): Agents that run software need a context graph of the app: every resource, what it connects to, the repository that deploys it and the team that owns it. How we built one on Durable Objects, keep it fresh across every provider without polling AWS, decide what counts as a change, and delete from it safely.
- [Cloudflare Durable Objects pending I/O keep-alive guide](https://polylane.com/learn/reliability/cloudflare-durable-objects-pending-i-o-keep-alive/): From 2026-10-01, pending I/O keeps Cloudflare Durable Objects in memory after the client leaves. What counts, the 15-minute limit, flags and billing.
- [Cloudflare K2: how to adopt serverless event streams](https://polylane.com/learn/reliability/cloudflare-k2-serverless-event-streams/): Cloudflare K2, announced 1 October 2026, is a serverless event stream on R2. What it is, how it differs from Queues, setup steps and beta limits.

Get started with one command: `curl -fsSL https://polylane.com/setup | bash` installs the CLI, connects your coding agents, and creates the account.
