Get started Dashboard

Reliability. How to set SLOs and error budgets, measure uptime honestly, and plan for the failures every production system has.

Reliability is how well a system keeps doing its job when parts of it fail, and every production system has parts that fail. It covers the targets you set, such as SLOs and error budgets, and the platform behaviour you design around: timeouts, retries, eviction, cold starts and limits.

These guides are for engineers building on serverless and managed platforms, where much of that behaviour belongs to the provider. They explain how specific primitives behave in production, such as Durable Objects staying in memory while I/O is pending, event streams on object storage and multi-node GPU clusters, and what that means for your design. Each guide covers setup, the limits to plan for, what it costs and how to tell from your own telemetry that it works as intended.

How Polylane works on these platforms: Cloudflare and Modal.

More guides

Nobody should be on-call. Polylane watches your infra, finds what broke, and writes the fix.

Get started for free