仪表盘

你的遥测早就知道哪里出了问题。 Polylane真的会去读它。

加入候补名单

它来盯着,你不必干瞪眼。 你的指标、日志和追踪全天候被读取:被判断,而不只是被存储。

Issues Critical latency degradation in checkout-edge worker Overview

Critical latency degradation in checkout-edge worker

Incident critical Polylane ·Detected 3 hours ago ·Last seen 4 minutes ago ·2 occurrences
OverviewInvestigationMetricsTimelineProperties
checkout-edge Cloudflare Worker

Critical latency degradation detected in checkout-edge worker: 18x+ P99 latency spikes sustained for 12 minutes

Causal metrics All metrics
Request Duration
414 ms ▲ 43.1σ
Wall Time
416 ms ▲ 43.0σ
CPU Time
14 ms ▲ 2.6σ
Blast radius
Search your cloud resources...
Graph Table
edge-gateway Cloudflare Worker ··· checkout-edge Cloudflare Worker ··· hd-prod Hyperdrive ··· payments-db PlanetScale ··· cart-svc Cloudflare Worker ···
Analysis Copy

Deploy 9f3c2a1 shrank the Hyperdrive pool hd-prod from 50 connections to 5. Under checkout load, requests queue on connection checkout and P99 rises 18× against the 30-minute baseline. Restoring the pool size restores latency.

Tags
service · checkout-edgeprovider · cloudflaredeploy · 9f3c2a1signal · wall_time_p99

它问的是你的代码能回答的问题。 监控由你代码中声明的日志行、span和指标构建而成。

Clouds prod-cloudflare Key queries

Key queries

31 active · 2 telemetry gaps

Generated from your telemetry and your connected code, judged on observed data: every candidate ran against a day of real data before it was kept.

Did any payment capture fail? code answered
logs: "payment capture failed for order {orderId}"
declared in apps/api/src/payments.ts
Is the checkout webhook retrying more than usual? telemetry answered
metric: webhook.delivery.retries · rate over 5m
baseline 0.4/min over the observed day
Is the nightly reconciliation completing? code answered
logs: "reconciliation complete" · count over 24h
quiet all day: healthy. Kept because the code provably emits it
Are D1 writes hitting rate limits? telemetry unanswerable
no series found for this question
kept as a telemetry gap
Re-confirmed on every discovery run: one thin run cannot wipe a good set.

缺口用代码补上。 Polylane替你写好缺失的埋点。

Clouds prod-cloudflare Key queries

Key queries

31 active · 2 telemetry gaps

Generated from your telemetry and your connected code, judged on observed data: every candidate ran against a day of real data before it was kept.

Did any payment capture fail? code answered
logs: "payment capture failed for order {orderId}"
declared in apps/api/src/payments.ts
Is the checkout webhook retrying more than usual? telemetry answered
metric: webhook.delivery.retries · rate over 5m
baseline 0.4/min over the observed day
Is the nightly reconciliation completing? code answered
logs: "reconciliation complete" · count over 24h
quiet all day: healthy. Kept because the code provably emits it
Are D1 writes hitting rate limits? telemetry unanswerable
no series found for this question
kept as a telemetry gap
Re-confirmed on every discovery run: one thin run cannot wipe a good set.

Polylane如何工作。 它学习你的系统、监视它、调查,然后行动。

  • 它先学习你的系统

    上下文图映射你的云、仓库和可观测性提供商中的每一个资源和依赖。智能体基于真实拓扑推理,而不是猜测。

  • 无阈值检测

    每个提供商都有内置检查,再加上由你自己保存的查询和仪表盘生成的检查。统计分析和智能体共同判断,改善永远不会触发问题。

  • 有据可查的调查

    每一条结论都链接回背后的查询、日志行或变更记录。没有证据的判定回落为无法确定。

  • 写入是争取来的,从不默认拥有

    账户以只读方式连接。回滚默认关闭、受限速并留有记录。代码变更走你正常的审查流程。

  • 它每周都更敏锐

    记忆、每日笔记和用真实数据反复确认的监控查询:七月的调查能从六月的调查中学习。

它接入你已经在运行的一切。 以只读方式连接,然后开始。

常见问题。

这会取代我的可观测性工具吗?

不会。Polylane不是仪表盘,也不是指标存储:它通过Datadog、Honeycomb、Axiom、Grafana Cloud和Sentry读取你仪表盘上绘制的那些序列,并做它们不做的那部分:判断、调查和行动。

如果我没有可观测性提供商怎么办?

云原生信号已经覆盖了大部分:Cloudflare上的Workers分析、AWS上的CloudWatch,以及其他平台上的对应物。连接Datadog、Honeycomb、Axiom、Grafana Cloud或Sentry后,这些序列会作为一等数据源加入同一张图。

什么是key queries?

每一个背后都有一条查询的监控问题,由你的遥测和已连接的代码生成,并在保留之前用一天的真实数据执行验证。账户目前还回答不了的问题会被有意保留,作为值得补上的遥测缺口。

埋点修复会改变我的技术栈吗?

不会。它们沿用你已经在用的配置,不添加新依赖:结构化事件、错误捕获,以及调查真正需要的属性。

AI监控会制造更多告警噪声吗?

默认判定是没有问题。检查知道哪个方向是变坏,改善永远不会触发,单纯的流量波动不会被当作故障。沉默是一项功能:Polylane只在真正出问题时开口。

更多使用场景

别再盯着仪表盘了。 开始阅读答案。