Cloudflare Durable Objects pending I/O keep-alive guide
Explore with AI
Since 1 October 2026, a Durable Object with no connected client stays in memory while it has pending service binding requests, RPC or fetch() calls to other Durable Objects, container.monitor() calls, waitUntil() promises or timers. This is the default for Workers with a compatibility date of 2026-10-01 or later; older Workers can add the durable_object_io_tasks_prevent_eviction flag, and durable_object_io_tasks_do_not_prevent_eviction turns it off. Each pending operation holds off idle eviction for up to 15 minutes from when it starts, and you pay duration charges the whole time it does.
On this page
- What pending I/O keep-alive means for a Durable Object
- Why an idle object could lose its work before 1 October 2026
- The operations that now keep an object in memory
- Fifteen minutes per operation, counted from when it starts
- Turning keep-alive on with a compatibility date or a flag
- Handing an agent job to the object and letting the client leave
- Restarts that keep-alive does not prevent
- What a held object costs in duration charges
- Where pending I/O keep-alive catches teams out
- A checklist before you bump the compatibility date
- Checking the Wrangler change against production with Polylane
- Common questions
Cloudflare announced on 1 October 2026 that pending I/O now keeps a Durable Object in memory after its client has gone. Before this, work an object started for a client could be lost once that client disconnected. With no client connected, Cloudflare could shut down an idle object while a service binding call, an RPC to another Durable Object, a waitUntil() promise or a timer was still pending. Those operations now hold off idle shutdown, each one for up to 15 minutes.
Agents gain the most from this. An agent can accept a job, reply to the caller at once, and carry on calling tools through service bindings, coordinating with other Durable Objects, or waiting on a container process. To turn it on, use a compatibility date of 2026-10-01 or later, or add a flag on older Workers. You pay duration charges for every second it keeps the object in memory.
What pending I/O keep-alive means for a Durable Object
A Durable Object stays active while it handles a request from a connected client, and that part hasn’t changed. The new behaviour covers the time after the client has gone, when the object still has work in flight. Pending I/O is any operation the object has started and is still waiting on. Under the new rules, each of those operations counts as a reason to keep the object in memory. Open outbound connections already worked this way.
The announcement names the case it was built for: an agent that keeps working on a submitted job after its client disconnects. It also fits job runners that call internal Workers, objects that coordinate a group of other Durable Objects, and objects that watch a container with this.ctx.container.monitor().
Why an idle object could lose its work before 1 October 2026
The Durable Objects lifecycle has five states: active in memory, idle in memory and hibernateable, idle in memory and non-hibernateable, hibernated, and inactive. An object can only hibernate when nothing is pending: no setTimeout or setInterval callbacks, no unfinished I/O or waitUntil() promise, no open outbound connection, no standard WebSocket API in use, and no request still being processed. When all of that holds, the object hibernates after 10 seconds without a request or event.
An object with pending I/O fails that test, so it sits in the idle, non-hibernateable state. Under the older rules, an object in that state was evicted after 70 to 140 seconds without incoming requests or events, even while a service binding call or RPC was still waiting. The unfinished work stopped with it.
Outbound connections were fixed first. A changelog entry on 19 June 2026 made TCP sockets opened with connect() and outbound WebSockets keep a Durable Object alive. The October announcement notes that outbound fetch() requests to external services already keep Durable Objects running. The 1 October change gives internal calls and scheduled work the same protection.
The operations that now keep an object in memory
New with the 1 October 2026 change:
- Pending service binding requests, while they wait for a response
- Calls to another Durable Object through RPC or
fetch() this.ctx.container.monitor()- Promises passed to
this.ctx.waitUntil() - Pending
setTimeout()andsetInterval()timers
Already protected before this change:
- Outbound
fetch()requests to external services - TCP sockets
- Outbound WebSockets
The waitUntil() reference states that a promise passed to it prevents eviction until it settles or for up to 15 minutes, whichever comes first. It does not change when the request or RPC that started it completes. Your handler can return while the promise keeps the object alive.
Fifteen minutes per operation, counted from when it starts
The limit is easy to misread, so here are the rules as Cloudflare states them:
- Each operation prevents eviction until it completes or until 15 minutes after it started, whichever comes first.
- Operations started together do not combine their windows. Three calls started at the same moment all stop protecting the object at the same moment.
- Starting another operation later can extend the object’s time in memory, because the new operation brings its own window.
- The limit applies to each operation. Cloudflare does not apply it to the total time an object spends in memory.
- Once no operation is preventing eviction, the standard eviction rules apply again.
Here is how that plays out for an agent job that runs for longer than 15 minutes:
min 0 client calls submit(); run() starts inside waitUntil() window A: 0 to 15
min 1 client disconnects
min 2 service binding call to the tools Worker starts window B: 2 to 17, ends at 3 when it returns
min 14 RPC to a reviewer Durable Object starts window C: 14 to 29
min 15 window A lapses; the run() promise is still pending
min 16 window C still protects the object
min 22 last tool call returns, job writes its final checkpoint
The one waitUntil() wrapped around the whole job protects only its first 15 minutes. After that, the job stays in memory because it keeps starting fresh operations. If a job sits on a single call that takes longer than 15 minutes, or goes quiet after its windows have lapsed, the normal eviction rules can remove it. For outbound connections, the June changelog adds that the connection keeps working after 15 minutes. It just stops preventing eviction.
Turning keep-alive on with a compatibility date or a flag
- Open your Wrangler configuration and find
compatibility_date. - If you can move to 2026-10-01 or later, the new behaviour comes with the date:
{
"compatibility_date": "2026-10-01"
}
- A new date also turns on every other change dated up to it. Read the flags history on the compatibility flags page for everything between your old date and the new one. One example is that Node.js compatibility is on by default from 2026-08-04.
- If you can’t move the date yet, add the enabling flag on its own:
{
"compatibility_date": "2026-06-01",
"compatibility_flags": ["durable_object_io_tasks_prevent_eviction"]
}
compatibility_date = "2026-06-01"
compatibility_flags = [ "durable_object_io_tasks_prevent_eviction" ]
- To move to a new date but keep the old eviction behaviour, add the opt-out flag:
{
"compatibility_date": "2026-10-01",
"compatibility_flags": ["durable_object_io_tasks_do_not_prevent_eviction"]
}
- Deploy. You can also set flags in the Workers settings on the Cloudflare dashboard, or in the
metadatafield when you upload through the Workers Script or Versions API.
Handing an agent job to the object and letting the client leave
The pattern: store the job, start the work inside waitUntil(), return straight away, and save progress after every step. Each service binding call and each RPC to another Durable Object inside the loop is a pending operation with its own 15-minute window.
import { DurableObject } from "cloudflare:workers";
export class Agent extends DurableObject {
running = false;
// Called by a Worker over RPC. Returns once the job is stored.
async submit(job) {
await this.ctx.storage.put("job", job);
await this.ctx.storage.put("step", 0);
await this.ctx.storage.put("status", "running");
this.ctx.waitUntil(this.run());
return { accepted: true };
}
async run() {
if (this.running) return;
this.running = true;
try {
const job = await this.ctx.storage.get("job");
let step = (await this.ctx.storage.get("step")) ?? 0;
for (; step < job.steps.length; step++) {
// Service binding call: its own pending operation
const output = await this.env.TOOLS.runTool(job.steps[step]);
// RPC to another Durable Object: another pending operation
await this.env.REVIEWER.getByName(job.id).record(step, output);
// Checkpoint before the next step
await this.ctx.storage.put("step", step + 1);
}
await this.ctx.storage.put("status", "done");
} finally {
this.running = false;
}
}
async status() {
return {
status: await this.ctx.storage.get("status"),
step: await this.ctx.storage.get("step"),
};
}
}
The caller gets { accepted: true } back right away and can poll status() later from any client. Don’t await this.run() inside submit(). That keeps the caller waiting for the whole job, which defeats the point of handing it off.
Restarts that keep-alive does not prevent
Keep-alive only affects idle eviction. Cloudflare still shuts Durable Objects down for new deployments, Workers runtime updates and decisions about where to host an object. When that happens, in-flight HTTP and RPC requests may finish only if they don’t touch storage. A request that tries to access storage is stopped with an error. During runtime updates, in-flight requests get up to 30 seconds to complete.
Cloudflare provides no shutdown hooks, so the only protection is to write state as you go. I built and led the Workers observability team at Cloudflare, and my advice here is to make unfinished work visible: no hook runs when an object is removed, so a stopped job looks exactly like a running one unless it records its own progress.
An alarm makes a cheap watchdog. An alarm is an event, so it starts a fresh instance after a restart, and in-memory state such as running begins empty in that new instance:
async submit(job) {
// store the job as above, then arm the watchdog
await this.ctx.storage.setAlarm(Date.now() + 5 * 60 * 1000);
this.ctx.waitUntil(this.run());
return { accepted: true };
}
async alarm() {
const status = await this.ctx.storage.get("status");
if (status === "running") {
await this.ctx.storage.setAlarm(Date.now() + 5 * 60 * 1000);
if (!this.running) this.ctx.waitUntil(this.run());
}
}
The step after the last checkpoint can run twice after a restart, so make each step safe to repeat.
What a held object costs in duration charges
Durable Objects are billed for wall-clock duration while they are active, or idle in memory and unable to hibernate. An object held by pending I/O is in that second state, and the announcement says duration charges continue while an operation prevents eviction. Duration is billed on the 128 MB allocated to the object, whatever it actually uses. On the Workers Paid plan, 400,000 GB-s per month are included, then $12.50 per million GB-s. The Free plan includes 13,000 GB-s per day.
one object held for a full 15-minute window
900 s x 128 MB / 1 GB = 115.2 GB-s
1,000 objects each held that long = 115,200 GB-s
For one job this is small. It adds up when objects hold themselves open for no reason, such as a promise that never settles or a forgotten setInterval() heartbeat. The same pending timer also stops the object from hibernating, so an app built on the WebSocket Hibernation API pays for idle time it expected to get free.
Where pending I/O keep-alive catches teams out
- The date never moved. Code written for the new behaviour still loses work on a Worker dated before 2026-10-01 with no flag. Fix: check
compatibility_dateandcompatibility_flagsin every Worker that hosts the class. - One
waitUntil()around a long job. Its window starts when the promise starts, so it protects only the first 15 minutes. Fix: rely on the fresh operations inside the job, and checkpoint each step. - Parallel calls treated as extra time. Operations started together share one end point. Fix: work out timings from the latest operation start, without adding windows together.
- A single call slower than 15 minutes. Once its window lapses, the standard rules apply again. Fix: split the work, or poll in shorter calls.
- Stray timers and promises. They keep the object in memory and billed for up to 15 minutes each. Fix: clear intervals when the work ends, and give external waits a timeout.
- Keep-alive treated as durability. Deploys and runtime updates still restart objects, and storage access in an in-flight request fails during shutdown. Fix: persist progress and resume from storage.
- Instance limits. The June changelog notes that the per-account instance limits for Durable Objects still apply.
- Cost-sensitive objects that relied on early eviction. Fix: add
durable_object_io_tasks_do_not_prevent_evictionfor those Workers.
A checklist before you bump the compatibility date
- List every Durable Object class that starts work in
waitUntil(), timers, service binding calls, Durable Object RPC orcontainer.monitor(). - Read the compatibility flag changes between your current date and 2026-10-01.
- Persist each job and its step counter before starting the work.
- Make each step safe to repeat after a restart.
- Arm an alarm or another watchdog that resumes stalled jobs.
- Clear timers and add timeouts so nothing stays pending without a reason.
- Estimate duration charges for your longest jobs at 128 MB each.
- Decide which Workers, if any, need the opt-out flag.
Checking the Wrangler change against production with Polylane
Polylane links each repository to the Workers it deploys by reading manifest files such as Wrangler files, then reviews every pull request against those live resources and posts a pass or fail comment. It also raises an advisory when a Cloudflare Worker has logs disabled. Without logs, a background job that stopped halfway leaves no trace.
Running on Cloudflare? See how Polylane monitors Cloudflare in production.
Common questions.
Do I need to change my code to get pending I/O keep-alive?
Not if your Worker uses a compatibility date of 2026-10-01 or later, because that is now the default. On an earlier date, add durable_object_io_tasks_prevent_eviction to compatibility_flags in your Wrangler configuration. Your code still has to put the work inside a pending operation, such as a promise passed to this.ctx.waitUntil().
Is 15 minutes the maximum length of a background job?
No. The 15-minute limit applies to each pending operation, counted from when it starts, and Cloudflare does not apply it to the total time in memory. Starting another operation later can extend the object's life. One operation that waits longer than 15 minutes stops protecting the object, and then the standard eviction rules apply again.
Does keep-alive stop deploys from restarting my Durable Object?
No. New deployments, Workers runtime updates and hosting decisions still shut objects down. In-flight requests that access storage during shutdown are stopped with an error, and during runtime updates in-flight requests have up to 30 seconds to finish. Write progress to storage as you go, so a restarted object can resume.
Will pending I/O keep-alive increase my Cloudflare bill?
It can. Duration charges continue while an operation prevents eviction, billed on the 128 MB allocated to the object. The Workers Paid plan includes 400,000 GB-s per month, then charges $12.50 per million GB-s. A forgotten setInterval() or a promise that never settles keeps an object billed for up to 15 minutes per operation.
How do I opt out of the new eviction behaviour?
Add the durable_object_io_tasks_do_not_prevent_eviction compatibility flag to the Worker. This lets you move to a compatibility date of 2026-10-01 or later and keep the older behaviour, where an idle object with no connected client could be evicted while these operations were still pending.
Did outbound fetch and WebSockets change on 1 October 2026?
No. Outbound fetch() requests to external services, TCP sockets and outbound WebSockets already kept Durable Objects running. A changelog entry on 19 June 2026 covered TCP sockets and outbound WebSockets, with the same 15-minute limit per connection. The October change adds service binding requests, Durable Object RPC and fetch() calls, container.monitor(), waitUntil() promises and timers.
Does waitUntil() make the caller wait for the background work?
No. Cloudflare's reference says a promise passed to waitUntil() does not affect when a request or RPC completes. The handler can return to the client at once while the promise keeps the object in memory until it settles, or for up to 15 minutes.
Can a Durable Object hibernate while it has pending I/O?
No. Hibernation requires that no I/O or waitUntil() promise is unfinished, no timers are set and no outbound connection is open. An object with pending work stays idle and non-hibernateable, and it keeps incurring duration charges. Once nothing is pending, it can hibernate after 10 seconds of inactivity.
Sources
- Pending I/O operations allow Durable Objects to continue long-running work without a connected client (Cloudflare changelog)
- Lifecycle of a Durable Object (Cloudflare Docs)
- Durable Object State (Cloudflare Docs)
- Durable Objects pricing (Cloudflare Docs)
- Outbound connections keep Durable Objects alive (Cloudflare changelog)
- Compatibility flags (Cloudflare Workers Docs)
- Polylane documentation
- Polylane full content
Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.
Related
- How to Debug Cloudflare Workers Errors: Logs, Traces and Error Codes
Debug Cloudflare Workers errors step by step: read 1101 and 1102 codes, enable Workers Logs and source maps, use wrangler tail, DevTools and local traces.
- Cloudflare K2: how to adopt serverless event streams
Cloudflare K2, announced 1 October 2026, is a serverless event stream on R2. What it is, how it differs from Queues, setup steps and beta limits.
- Turn your app into a context graph
Agents that run software need a context graph of the app: every resource, what it connects to, the repository that deploys it and the team that owns it. How we built one on Durable Objects, keep it fresh across every provider without polling AWS, decide what counts as a change, and delete from it safely.
- Cloudflare Containers agent sandboxes: startup and setup
Cloudflare Containers now start agent sandboxes in a median 648 ms. Set up the durable_object policy, runtime images and snapshots, and know the limits.
- Fix “Durable Object reset because its code was updated”
Why a deploy makes Cloudflare Durable Objects throw “reset because its code was updated”, and how to retry safely, keep clients connected and lose no state.