What Is an AI SRE? How It Works and How to Evaluate One
Explore with AI
An AI SRE is an AI agent that connects to your observability tools, cloud accounts and code. It triages alerts, investigates incidents and proposes fixes without waiting for someone to prompt it. It uses a large language model with tool access to query logs, metrics, traces and deploy history, and it should show the evidence behind each conclusion while a person approves risky actions.
AI SRE, defined
An AI SRE (AI site reliability engineer) is software that does the investigative part of on-call work. It watches telemetry and triages alerts. It works out what changed, finds the likely root cause and proposes a fix. It works as an autonomous agent: it decides what to query and acts on its findings without a human typing a prompt.
The term is new and loose. Many products now carry the label, so the label alone tells you little. What matters is what the tool does when an alert fires at night.
How it works
Most AI SREs follow the same loop, whatever the vendor.
- Build context. The agent connects to your clouds, repositories, observability tools and chat. It maps services, their dependencies, recent deploys and past incidents, so it reasons about your actual system. Good tools ground the model with retrieval over your runbooks, post-mortems and service topology.
- Triage. When an alert fires, the agent reads the data around it and groups related alerts into one incident. Then it decides whether the problem is real. Most alerts need no action, so this step removes a lot of noise.
- Investigate. It forms several hypotheses, such as a bad deploy, a config change or a saturated dependency. It checks each one against logs, metrics, traces and code. Good tools test hypotheses in parallel and drop the ones the evidence rules out.
- Explain. It writes up the cause as a chain: the signal, the change that triggered it, the code path and the blast radius. Each link should cite a query, log line or commit.
- Remediate. It proposes a rollback, a config change or a code fix as a pull request. A person approves anything risky.
- Document and learn. It drafts the timeline and post-mortem. It also saves confirmed findings so the next investigation starts faster.
AI SRE vs AIOps, copilots and runbooks
These tools overlap, so it helps to be precise.
- Runbook automation runs scripts someone wrote in advance. It handles only the failures you predicted.
- AIOps groups and deduplicates alerts with statistics. It tells you signals are related and leaves the cause to you.
- Copilots answer questions when you ask them.
- AI SREs start work when the incident starts and carry it from triage to a proposed fix.
Also watch for AI assistants built into a single observability platform. They can only reason over the data that platform holds, and many incidents cross a cloud, a database and a deploy.
Levels of autonomy
Autonomy comes in four levels, and most teams move up them one action type at a time:
- Observe and report: the agent investigates every alert and posts findings with evidence.
- Advise: it ranks remediations by confidence, and you approve or reject each one.
- Act with approval: it prepares the fix and runs it after one human approval.
- Autonomous for known patterns: it resolves well-understood, low-risk incidents on its own, within the policy you set.
Start at level one. Raise autonomy for an action type once the agent has been right on it many times. Keep anything that touches customer data behind approval.
How to evaluate an AI SRE
Run a trial on your own stack with these steps:
- Pick a few resolved incidents where you already know the root cause. Ask the tool about the same time window and compare its answer with yours.
- Check every claim for a citation. Click through to the query, log line or commit. A step with no evidence should be marked as unknown.
- Connect with read-only credentials first. Confirm that write actions pause for approval and that code changes arrive as pull requests your reviewers and CI gate.
- Count the noise. Track how many alerts it dismisses and spot-check those dismissals.
- Test a cross-service incident, such as a deploy that exhausts a database connection pool. See whether it follows the path from the symptom to the change.
- Measure the time from alert to a root cause you trust, and compare it with your current baseline.
Get your stack ready
An agent inherits the quality of your telemetry. You can improve results today with the tools you already have:
- Emit structured logs with a service name and request ID on every event.
- Record deploys with commit SHAs, so changes line up with metric shifts.
- Keep infrastructure-as-code and deploy manifests in the repository, so a tool can link code to the resources it ships to.
- Write down service ownership, for example in a CODEOWNERS file, and keep runbooks close to the code.
- Save the dashboards and queries your team trusts. Many tools reuse them as checks.
Common mistakes
- Trusting a confident answer without reading the evidence. Language models can produce plausible wrong root causes. The safeguard is a glass box: every suggestion links to the data behind it.
- Granting write access on day one. Earn autonomy per action type.
- Judging each step in isolation. Evaluate the whole run on a real incident, from alert to fix.
- Expecting it to replace engineers. It takes on the toil. People keep design, risk calls and new failure modes with no history to learn from.
Doing this with Polylane
Polylane connects your clouds, repositories and observability tools, raises issues from its checks and forwarded alerts, and gives each issue to one agent that reaches a verdict, writes a causal chain with every link backed by evidence or marked unknown, and opens the fix as a pull request. Cloud accounts start read-only, write calls wait for your confirmation in the thread, and the merge stays with you.
Common questions.
Will an AI SRE replace my SRE team?
No. It takes on the toil of triage, investigation and documentation, while people keep architecture, risk decisions and new failures with no precedent.
Is it safe to give an AI SRE access to production?
Start with read-only access and require human approval for any write. Let the agent propose fixes for engineers to approve, and send code changes through your normal review.
How is an AI SRE different from AIOps?
AIOps groups and deduplicates alerts with statistical models. An AI SRE uses a language model with tool access to investigate, explain the cause with evidence and propose a fix.
How do I evaluate an AI SRE?
Trial it on incidents you have already resolved: ask about the same time window and compare its answer with the root cause you found. Check that every step cites a query, log line or commit, connect it read-only first, and measure the time from alert to a root cause you trust.
What should I have in place before adopting an AI SRE?
Good telemetry, because the agent inherits its quality. Emit structured logs with a service name and request ID, record deploys with commit SHAs, keep infrastructure-as-code and deploy manifests in the repository, and write down who owns each service.
Sources
Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.
Related
- Workers Issues: route Cloudflare errors to coding agents
Cloudflare's new Workers Issues groups exceptions and 5xx errors and sends them to Claude Code, Cursor, Devin or a webhook. Setup, automations and limits.
- Fix Cloudflare “Error 1102: Worker exceeded resource limits”
Why a Cloudflare Worker returns Error 1102, how to tell a CPU time overrun from a 128 MB memory overrun, and the code and Wrangler changes that fix each.
- Fix Vercel “FUNCTION_INVOCATION_FAILED” 500 errors
Why a Vercel Function returns 500 FUNCTION_INVOCATION_FAILED, how to find the crash in runtime logs, and how to fix throws, missing env vars, files and memory.
- Fix Vercel “MIDDLEWARE_INVOCATION_FAILED” errors
Why Vercel returns 500 MIDDLEWARE_INVOCATION_FAILED or EDGE_FUNCTION_INVOCATION_FAILED, how to find the failing middleware in logs, and how to fix it.
- Cloudflare Workers error 1101: causes and how to fix it
Error 1101 means your Cloudflare Worker threw an uncaught JavaScript exception. Find the exception in logs, match it to its cause, fix it and roll back fast.