Get started Dashboard
AWS Render Kubernetes Axiom Cloudflare

Find the failures that never alerted

Sweeps AWS, Render, Kubernetes and Axiom for the failures no alert rule covers: dead crons, growing queues, errors nobody sees.

I want to do this: Find things that look wrong but never alerted: error spikes, dead crons, growing queues.

## Setup (skip if Polylane is already set up)

Read and follow https://polylane.com/auth.md for non-interactive signup and setup. Start by checking whether I am already signed in; reuse my account and workspace.

If the CLI is missing, bootstrap it without starting the interactive wizard:

curl -fsSL 'https://polylane.com/setup?ref=prompts' | bash -s -- --install-only

Then follow the guide through email verification, workspace selection, source connections, and MCP authentication. Ask me for an email code or OAuth consent only when needed. Verify each step; report pending setup instead of claiming success from installation alone.

## How to work

Over MCP: searchTools lists what this workspace exposes, with each tool's schema; call it first. runTool runs one tool, runCode chains several in one call and returns just the answer. search and execute cover the full Polylane REST API: threads, issues, investigations, autofixes, memories.
From the terminal: the polylane CLI wraps the same API, with structured output and non-interactive flags everywhere.
Reads always work. Write tools appear only if I have opted in, and every write is screened.

## Task: Find the failures that never alerted

Steps:
1. Sweep scheduled jobs and workers across providers for anything that stopped succeeding
2. Check queue depth and message age against their own history
3. Scan error rates per service for sustained drift, not just spikes
4. Cross-check each finding against open issues to skip the known ones
5. Deliver the list of quiet failures, each with the series that exposes it

Ground every claim in data you actually pulled: the query, the log line, the change record. If the data is inconclusive, say so. Ask me before anything that writes.

Alert rules only catch what someone predicted.

The cron that stopped in July, the queue growing one percent a day, the error rate that doubled from a tiny base: none of them page anyone, because nobody wrote that rule.

  • The dead cron discovered by the customer it served
  • Queues that grow quietly until they don't
  • Coverage defined by which rules someone remembered to write

One prompt, this much work. Every step on your real data.

  1. 1 Sweep scheduled jobs and workers across providers for anything that stopped succeeding AWS Render Kubernetes
  2. 2 Check queue depth and message age against their own history AWS Cloudflare
  3. 3 Scan error rates per service for sustained drift, not just spikes Axiom
  4. 4 Cross-check each finding against open issues to skip the known ones
  5. 5 Deliver the list of quiet failures, each with the series that exposes it

The failures nobody was watching, listed

A list of things going wrong that no rule was watching, each with the evidence. The next incident gets caught while it's still a curiosity.

More prompts for DevOps

Stop doing this by hand. Paste it, and your agent does the rest.