別再寫警示規則了。 代理讀你的日誌,然後做決定。
加入候補名單 Issues Critical latency degradation in checkout-edge worker Overview
Critical latency degradation in checkout-edge worker
Incident critical Polylane ·Detected 3 hours ago ·Last seen 4 minutes ago ·2 occurrences
OverviewInvestigationMetricsTimelineProperties
Critical latency degradation detected in checkout-edge worker: 18x+ P99 latency spikes sustained for 12 minutes
Causal metrics All metrics
Request Duration
414 ms ▲ 43.1σ
Wall Time
416 ms ▲ 43.0σ
CPU Time
14 ms ▲ 2.6σ
Blast radius
edge-gateway Cloudflare Worker ···
checkout-edge Cloudflare Worker ···
hd-prod Hyperdrive ···
payments-db PlanetScale ···
cart-svc Cloudflare Worker ···
Search your cloud resources...
5 resources Graph Table
Analysis Copy
Deploy 9f3c2a1 shrank the Hyperdrive pool
hd-prod from 50 connections to 5. Under checkout load,
requests queue on connection checkout and P99 rises 18× against the 30-minute baseline. Restoring the pool size restores latency.
Tags
service · checkout-edgeprovider · cloudflaredeploy · 9f3c2a1signal · wall_time_p99
一次尖峰不等於一次事件。 Polylane 分得出差別。
-
你的圖表變成檢查
你團隊已經在畫的圖表都會被監看:儲存一個查詢或建一個儀表板,Polylane 就會接手。
-
懂你流量的基準線
週二下午兩點的正常,不是週日凌晨兩點的正常。Polylane 對照的是你的節奏,不是一條平線。
-
小波動會被忽略
變化必須夠大、夠持久,而且往變糟的方向走。孤立的尖峰會被排除。
-
由代理做出判斷
它看的是全貌,包括最近改了什麼。預設的答案是一切正常。
-
附上證據
當真的出問題時,你拿到的是發現背後確切的查詢,而不是一張圖加一句祝你好運。
有些問題從來不會出現尖峰。 Polylane 也會為那些問題去讀。
一連串持續的崩潰在每一張圖表上都看起來是平的,所以代理也會自己讀日誌。而有些狀態在任何基準線下都是錯的:一個 CrashLoopBackOff 不需要判斷,一個簡單的閾值就能立刻抓到。