Kill the bug that survived the backlog
Runs the investigation nobody had time for, across every Cloudflare, Datadog, Sentry and Axiom signal the flakiness ever touched.
I want to do this: This flaky 502 has been in the backlog for a month. Find the actual cause. ## Setup (skip if Polylane is already set up) Read and follow https://polylane.com/auth.md for non-interactive signup and setup. Start by checking whether I am already signed in; reuse my account and workspace. If the CLI is missing, bootstrap it without starting the interactive wizard: curl -fsSL 'https://polylane.com/setup?ref=prompts' | bash -s -- --install-only Then follow the guide through email verification, workspace selection, source connections, and MCP authentication. Ask me for an email code or OAuth consent only when needed. Verify each step; report pending setup instead of claiming success from installation alone. ## How to work Over MCP: searchTools lists what this workspace exposes, with each tool's schema; call it first. runTool runs one tool, runCode chains several in one call and returns just the answer. search and execute cover the full Polylane REST API: threads, issues, investigations, autofixes, memories. From the terminal: the polylane CLI wraps the same API, with structured output and non-interactive flags everywhere. Reads always work. Write tools appear only if I have opted in, and every write is screened. ## Task: Kill the bug that survived the backlog Steps: 1. Gather every occurrence from the telemetry: full history, not last week's 2. Find what the occurrences share: route, region, time, payload shape, upstream 3. Correlate the pattern with changes and infrastructure events over the same period 4. Test the surviving theory against the accumulated evidence 5. Deliver the cause with the pattern analysis, and the fix drafted if it's code Ground every claim in data you actually pulled: the query, the log line, the change record. If the data is inconclusive, say so. Ask me before anything that writes.
Flaky bugs win by being boring.
It fires once a day, hurts one user at a time, and never justifies a sprint. So it sits in the backlog collecting duplicates while its trail goes cold every single night.
- Duplicate reports stapled to a ticket nobody reads
- Each investigation starting from zero, months apart
- "Low priority" bugs that quietly churn customers
One prompt, this much work. Every step on your real data.
- 1 Gather every occurrence from the telemetry: full history, not last week's
- 2 Find what the occurrences share: route, region, time, payload shape, upstream
- 3 Correlate the pattern with changes and infrastructure events over the same period
- 4 Test the surviving theory against the accumulated evidence
- 5 Deliver the cause with the pattern analysis, and the fix drafted if it's code
A month of flakiness, one afternoon of analysis
The whole history analysed at once instead of one cold occurrence at a time. The pattern that no single incident showed becomes obvious, and the ticket finally closes.
Ticket to fix
“Users report intermittent timeouts on search. Find the cause and draft the fix.” Confirm the bug
“Can you confirm the bug in ticket #513 from the logs, and which release introduced it?” Answer support
“What should support tell customers about yesterday's checkout errors? Keep it accurate.”