Get started Dashboard

The ticket becomes a pull request. Polylane does the digging in between.

Polylane turns a bug report into a proven cause and a fix to review.

Get started for free

Paste the ticket. That's the whole handover. No repro steps, no triage meeting, no context to assemble.

Threads Critical latency degradation in checkout-edge worker Share

What I did

I queued an autofix that restores the Hyperdrive pool hd-prod in the coreplane/checkout-edge repo with:

  • max_connections restored from 5 to 50 on the pool binding
  • a regression test that fails when checkout queues on the pool
  • wrangler.jsonc configured to deploy as checkout-edge

This gives Boris a PR with deployable code. Once merged, running wrangler deploy will replace the degraded worker with the fixed version.

The root cause report has been updated with the occurrence #2 evidence: Root Cause: checkout-edge Latency Degradation (Recurring).

Wall Time
checkout-edge worker · last hour
Export chart
Wall Time P99
28,381ms
Wall Time P95
2,696ms
Wall Time P50
118ms
010k20k30k
18:0018:2519:00
Worked for 4m 12s

AI can make mistakes. Please double check cited sources.

Dig deeper...

It proves the cause before it writes the fix. Each step names the evidence under it.

Critical latency degradation in checkout-edge worker Timeline
  1. 02:02:31 Fix run started

    One agent, from the signal to the pull request.

  2. 02:16:41 Causal chain
    Signal (metric)

    workers.invocation.p99_latency · checkout-edge

    Surfacing site

    checkout-edge at coreplane/checkout-edge:src/hyperdrive/pool.ts#withConnection

    Mechanism

    withConnection waits for a free slot before it runs the query, so the P99 carries pool queue time, not database time

    Producer

    hd-prod, at coreplane/checkout-edge:wrangler.jsonc#hyperdrive

    • Producer evidence
    • queue-depth histogram grouped by colo: p99 wait 1.31s at 5 connections in flight, 0.04s below
    • sandboxGrep: max_connections is 5 on the hd-prod binding; the worker opens up to 50
    • change record 9f3c2a1: max_connections 50 → 5
    Trigger

    deploy 9f3c2a1 took max_connections from 50 to 5, 18 minutes before the first breach

    What happens to the failed unit today

    unknown: whether the storefront retries the confirm: coreplane/storefront-web is not a connected repository

    Cadence check

    one queued request per checkout past the fifth in flight matches the 118 breaches and the gaps between evening peaks

    Blast radius

    2 other resource(s), 0 other tenant(s); data at risk: none named

    Every link rests on named evidence, or is written as unknown: <what would establish it>. Counts, not adjectives. If the cause isn't proven, no fix ships.

  3. 02:24:09 Pull request #412

    Restore Hyperdrive pool size in checkout-edge

The fix arrives ready to review. Root cause, validation, and the diff included.

github.com/coreplane/payments-api/pull/491

Cap retries on the checkout webhook worker #491

polylane
Open polylane wants to merge 1 commit into main from polylane/autofix/chat/k3x9f2-4e7d21a
Conversation 1 Commits 1 Checks 1 Files changed 2
polylane bot commented 6 minutes ago ···

Retries on the checkout webhook worker were unbounded: a failing delivery re-queued itself forever and amplified load on payments-api. This caps delivery at 5 attempts with exponential backoff and dead-letters the payload after the last one.

What changed

worker/deliver.ts gains MAX_DELIVERY_ATTEMPTS = 5 and backoff between attempts; exhausted payloads land in checkout-webhooks-dlq instead of re-queueing.

Validation

npm test: 214 passed. A forced failing delivery stopped after 5 attempts and appeared in the dead-letter queue.

Root cause · Why it's safe · Out of scope
polylane added commit 4e7d21a Verified
Review required At least 1 approving review is required
ci / test Successful in 3m 12s Details
Review required Waiting on your review before it can merge
Merging is blocked

Peace of mind. Nothing happens behind your back.

Agents should act on your behalf without putting production at risk. Polylane is built so you can leave it alone.

Read-only from the start

Every account connects read-only: AWS through a scoped role, most providers with an API token. Write access is a separate, per-account decision.

Agent
AWS Cloudflare
query metrics delete stack read logs purge cache

Writes are earned

A write pauses the run and shows you the exact call before anything happens, unless you allowed it ahead of time. Nothing reaches production just because an agent decided it should.

write request

PATCH /workers/checkout-edge

reason: restore the Hyperdrive pool to 50

Approve Reject with a note
waiting on you Cloudflare

Evidence for everything

Every claim links to the query, log line or deploy behind it. Every run is a transcript you can watch, interrupt and share.

thread transcript

● queryLogs → 4,112 rows cited

● deployHistory → 12 deploys cited

cause: deploy 9f3c2a1 evidence

These prompts already work. Copy one, swap your own service name in, and run it.

Give it to your coding agent and let it cook.

Every prompt

Swap the agent. Keep everything.

Your connections, context graph and memories live with Polylane, not with the agent. When a better coding agent ships, point it at the same MCP server and everything comes along.

  • Connections and scopes live in the workspace: connect once, every agent benefits
  • Works with Claude Code, Cursor, Codex, OpenCode, VS Code, Pi, and whatever ships next
  • Memories, notes and investigations persist: switching agents loses nothing
Claude Code Cursor Codex OpenCode VS Code Pi
Claude CodeCursorCodexOpenCodeVS CodePi
aws cloudflare datadog github slack sentry

Polylane runs the whole loop. You step in when it needs you.

  • It learns your system first

    The context graph connects every resource and dependency across your clouds, repos and observability providers, so Polylane reasons over your real system instead of guessing.

  • Detection without thresholds

    Built-in checks for every provider, plus checks built from your own saved queries and dashboards. Statistics and an agent decide together, and a metric getting better never raises an issue.

  • Investigations that cite their sources

    Every claim links back to the query, log line or deploy behind it. Without evidence, the verdict is inconclusive.

  • Writes are earned, never assumed

    Accounts connect read-only. Write actions pause for your approval, with the request and reason on screen, unless you've allowed it ahead of time. Code changes go through your normal review.

  • It gets sharper every week

    It remembers what it confirms, keeps daily notes, and re-checks its monitoring queries against real data, so every incident makes the next one faster.

It plugs into what you already run. Connect read-only and start.

Common questions.

How do I hand Polylane a ticket?

Describe the bug in a thread, over the CLI, or from your editor through MCP. Polylane already has the graph, the telemetry, the recent changes and the code.

What if it can't find the cause?

It says so. An investigation ends resolved, diagnosed or inconclusive, and anything that needs a human decision becomes a single escalation.

Does it learn from past tickets?

Yes. What it confirms on one ticket is remembered and used on the next, and daily notes keep a record of what happened.

Who writes the fix?

Polylane's own executor by default, or connect Devin, Cursor, Factory or Conductor and every autofix routes through it. Either way, you review it before it merges.

Which repositories does it work with?

GitHub repositories, connected through the GitHub app. The fix lands on a branch in the repo behind the affected service, with the investigation linked from the change.

More use cases

Tickets in. Pull requests out. The backlog finally moves on its own.