Get started Dashboard

It's 2:14am and checkout is failing. By 2:15 the investigation is running.

Polylane triages the alert, works the incident, and shows its work.

Get started for free

An alert fires. Polylane picks it up. Real or noise, usually answered inside a minute.

Issues Elevated error rate on payments-api Timeline

Elevated error rate on payments-api

High Confirmed ·Detected from a Datadog alert
OverviewMetricsLogsTracesTimelinePropertiesInvestigation
  1. 02:14:07 Alert received

    Datadog · payments-api error rate · occurrence #1

  2. 02:14:58 Triaged as an incident

    Severity high. Error rate 41× the 30-minute baseline, clustered on POST /v1/charges; deploy f3a9c12 landed 18 minutes before the firing.

  3. 02:15:02 Fix run started

    The same agent that triaged it keeps working it.

  4. 02:31:44 Alert fired again (occurrence #2)

    Folded into this issue after 16 quiet minutes, and handed to the investigation already running.

  5. 03:06:12 Alert recovered

    Provider recovery received. Issue resolved.

It digs until it can prove it. One agent, one transcript, evidence under every claim.

Threads Critical latency degradation in checkout-edge worker Share

What I did

I queued an autofix that restores the Hyperdrive pool hd-prod in the coreplane/checkout-edge repo with:

  • max_connections restored from 5 to 50 on the pool binding
  • a regression test that fails when checkout queues on the pool
  • wrangler.jsonc configured to deploy as checkout-edge

This gives Boris a PR with deployable code. Once merged, running wrangler deploy will replace the degraded worker with the fixed version.

The root cause report has been updated with the occurrence #2 evidence: Root Cause: checkout-edge Latency Degradation (Recurring).

Wall Time
checkout-edge worker · last hour
Export chart
Wall Time P99
28,381ms
Wall Time P95
2,696ms
Wall Time P50
118ms
010k20k30k
18:0018:2519:00
Worked for 4m 12s

AI can make mistakes. Please double check cited sources.

Dig deeper...

The incident ends with a fix, not a follow-up meeting. The cause found and the fix written before anyone asks.

github.com/coreplane/payments-api/pull/491

Cap retries on the checkout webhook worker #491

polylane
Open polylane wants to merge 1 commit into main from polylane/autofix/chat/k3x9f2-4e7d21a
Conversation 1 Commits 1 Checks 1 Files changed 2
polylane bot commented 6 minutes ago ···

Retries on the checkout webhook worker were unbounded: a failing delivery re-queued itself forever and amplified load on payments-api. This caps delivery at 5 attempts with exponential backoff and dead-letters the payload after the last one.

What changed

worker/deliver.ts gains MAX_DELIVERY_ATTEMPTS = 5 and backoff between attempts; exhausted payloads land in checkout-webhooks-dlq instead of re-queueing.

Validation

npm test: 214 passed. A forced failing delivery stopped after 5 attempts and appeared in the dead-letter queue.

Root cause · Why it's safe · Out of scope
polylane added commit 4e7d21a Verified
Review required At least 1 approving review is required
ci / test Successful in 3m 12s Details
Review required Waiting on your review before it can merge
Merging is blocked

Peace of mind. Nothing happens behind your back.

Agents should act on your behalf without putting production at risk. Polylane is built so you can leave it alone.

Read-only from the start

Every account connects read-only: AWS through a scoped role, most providers with an API token. Write access is a separate, per-account decision.

Agent
AWS Cloudflare
query metrics delete stack read logs purge cache

Writes are earned

A write pauses the run and shows you the exact call before anything happens, unless you allowed it ahead of time. Nothing reaches production just because an agent decided it should.

write request

PATCH /workers/checkout-edge

reason: restore the Hyperdrive pool to 50

Approve Reject with a note
waiting on you Cloudflare

Evidence for everything

Every claim links to the query, log line or deploy behind it. Every run is a transcript you can watch, interrupt and share.

thread transcript

● queryLogs → 4,112 rows cited

● deployHistory → 12 deploys cited

cause: deploy 9f3c2a1 evidence

These prompts already work. Copy one, swap your own service name in, and run it.

Give it to your coding agent and let it cook.

Every prompt

Swap the agent. Keep everything.

Your connections, context graph and memories live with Polylane, not with the agent. When a better coding agent ships, point it at the same MCP server and everything comes along.

  • Connections and scopes live in the workspace: connect once, every agent benefits
  • Works with Claude Code, Cursor, Codex, OpenCode, VS Code, Pi, and whatever ships next
  • Memories, notes and investigations persist: switching agents loses nothing
Claude Code Cursor Codex OpenCode VS Code Pi
Claude CodeCursorCodexOpenCodeVS CodePi
aws cloudflare datadog github slack sentry

Polylane runs the whole loop. You step in when it needs you.

  • It learns your system first

    The context graph connects every resource and dependency across your clouds, repos and observability providers, so Polylane reasons over your real system instead of guessing.

  • Detection without thresholds

    Built-in checks for every provider, plus checks built from your own saved queries and dashboards. Statistics and an agent decide together, and a metric getting better never raises an issue.

  • Investigations that cite their sources

    Every claim links back to the query, log line or deploy behind it. Without evidence, the verdict is inconclusive.

  • Writes are earned, never assumed

    Accounts connect read-only. Write actions pause for your approval, with the request and reason on screen, unless you've allowed it ahead of time. Code changes go through your normal review.

  • It gets sharper every week

    It remembers what it confirms, keeps daily notes, and re-checks its monitoring queries against real data, so every incident makes the next one faster.

It plugs into what you already run. Connect read-only and start.

Common questions.

Does Polylane page me?

No, it isn't a pager. It tells you by email, Slack and the console, only for critical and high-severity issues it found itself, and each one arrives already investigated.

What happens when an issue is detected at night?

Polylane confirms it's real and starts investigating on its own. By morning the issue has a verdict and the evidence behind it, and where autofix is on, the fix too.

Which alert sources does Polylane ingest?

Datadog, Sentry, Honeycomb, Grafana Cloud, Axiom, Logfire, Better Stack, OpenStatus, Amazon CloudWatch, Cloudflare, Vercel, Render, Railway, and Linear. Each source connects with its own scoped token, and every alert gets triaged.

What if Polylane calls a real incident noise?

The reasoning stays attached, and one click sends it back for investigation. If Polylane couldn't reach the data, it won't call it either way.

How many investigations will it run?

As many as your credits cover: no separate investigation quota. Set a minimum severity and anything below it waits for someone to press Start investigation.

More use cases

Hand the nights to the agents. Keep the mornings.