แดชบอร์ด

telemetry ของคุณรู้อยู่แล้วว่าอะไรผิดปกติ Polylane อ่านมันจริง ๆ

ลงชื่อเข้า waitlist

มันเฝ้าดูเพื่อที่คุณจะไม่ต้องจ้อง metric, log และ trace ของคุณ ถูกอ่านตลอด 24 ชั่วโมง: ตัดสิน ไม่ใช่แค่เก็บ

Issues Critical latency degradation in checkout-edge worker Overview

Critical latency degradation in checkout-edge worker

Incident critical Polylane ·Detected 3 hours ago ·Last seen 4 minutes ago ·2 occurrences
OverviewInvestigationMetricsTimelineProperties
checkout-edge Cloudflare Worker

Critical latency degradation detected in checkout-edge worker: 18x+ P99 latency spikes sustained for 12 minutes

Causal metrics All metrics
Request Duration
414 ms ▲ 43.1σ
Wall Time
416 ms ▲ 43.0σ
CPU Time
14 ms ▲ 2.6σ
Blast radius
Search your cloud resources...
Graph Table
edge-gateway Cloudflare Worker ··· checkout-edge Cloudflare Worker ··· hd-prod Hyperdrive ··· payments-db PlanetScale ··· cart-svc Cloudflare Worker ···
Analysis Copy

Deploy 9f3c2a1 shrank the Hyperdrive pool hd-prod from 50 connections to 5. Under checkout load, requests queue on connection checkout and P99 rises 18× against the 30-minute baseline. Restoring the pool size restores latency.

Tags
service · checkout-edgeprovider · cloudflaredeploy · 9f3c2a1signal · wall_time_p99

มันถามคำถามที่โค้ดของคุณตอบได้ การมอนิเตอร์ที่สร้างจากบรรทัด log, span และ metric ที่โค้ดของคุณประกาศไว้

Clouds prod-cloudflare Key queries

Key queries

31 active · 2 telemetry gaps

Generated from your telemetry and your connected code, judged on observed data: every candidate ran against a day of real data before it was kept.

Did any payment capture fail? code answered
logs: "payment capture failed for order {orderId}"
declared in apps/api/src/payments.ts
Is the checkout webhook retrying more than usual? telemetry answered
metric: webhook.delivery.retries · rate over 5m
baseline 0.4/min over the observed day
Is the nightly reconciliation completing? code answered
logs: "reconciliation complete" · count over 24h
quiet all day: healthy. Kept because the code provably emits it
Are D1 writes hitting rate limits? telemetry unanswerable
no series found for this question
kept as a telemetry gap
Re-confirmed on every discovery run: one thin run cannot wipe a good set.

ช่องว่างถูกปิดด้วยโค้ด Polylane เขียน instrumentation ที่ขาดหายให้คุณ

Clouds prod-cloudflare Key queries

Key queries

31 active · 2 telemetry gaps

Generated from your telemetry and your connected code, judged on observed data: every candidate ran against a day of real data before it was kept.

Did any payment capture fail? code answered
logs: "payment capture failed for order {orderId}"
declared in apps/api/src/payments.ts
Is the checkout webhook retrying more than usual? telemetry answered
metric: webhook.delivery.retries · rate over 5m
baseline 0.4/min over the observed day
Is the nightly reconciliation completing? code answered
logs: "reconciliation complete" · count over 24h
quiet all day: healthy. Kept because the code provably emits it
Are D1 writes hitting rate limits? telemetry unanswerable
no series found for this question
kept as a telemetry gap
Re-confirmed on every discovery run: one thin run cannot wipe a good set.

Polylane ทำงานอย่างไร มันเรียนรู้ระบบของคุณ เฝ้าดู สืบสวน และลงมือ

  • มันเรียนรู้ระบบของคุณก่อน

    context graph วาดแผนที่ทุกทรัพยากรและการพึ่งพาทั่วคลาวด์ repo และผู้ให้บริการ observability ของคุณ เอเจนต์ใช้เหตุผลบน topology จริง ไม่ใช่การเดา

  • การตรวจจับโดยไม่ใช้ threshold

    check ในตัวสำหรับทุกผู้ให้บริการ บวก check ที่สร้างจาก query และแดชบอร์ดที่คุณบันทึกไว้เอง การประมวลผลทางสถิติและเอเจนต์ตัดสินร่วมกัน และการเปลี่ยนแปลงในทางที่ดีขึ้นไม่มีทางเปิด issue

  • การสืบสวนที่แสดงหลักฐาน

    ทุกข้ออ้างลิงก์กลับไปยัง query, บรรทัด log หรือบันทึกการเปลี่ยนแปลงที่อยู่เบื้องหลัง คำตัดสินที่ไม่มีหลักฐานจะถอยกลับเป็นสรุปไม่ได้

  • การเขียนต้องได้รับสิทธิ์ ไม่ใช่ถือว่ามี

    บัญชีเชื่อมต่อแบบอ่านอย่างเดียว rollback ปิดอยู่โดยค่าเริ่มต้น จำกัดอัตรา และมีบันทึกไว้ การเปลี่ยนโค้ดผ่านการรีวิวตามปกติของคุณ

  • มันคมขึ้นทุกสัปดาห์

    memory โน้ตประจำวัน และ query สำหรับมอนิเตอร์ที่ยืนยันซ้ำกับข้อมูลจริง: การสืบสวนเดือนกรกฎาคมเรียนรู้จากเดือนมิถุนายน

มันเสียบเข้ากับสิ่งที่คุณรันอยู่แล้ว เชื่อมต่อแบบอ่านอย่างเดียวแล้วเริ่มได้เลย

คำถาม

นี่มาแทนเครื่องมือ observability ของฉันไหม

ไม่ Polylane ไม่ใช่แดชบอร์ดหรือที่เก็บ metric: มันอ่านชุดข้อมูลเดียวกับที่แดชบอร์ดของคุณพล็อต ผ่าน Datadog, Honeycomb, Axiom, Grafana Cloud และ Sentry แล้วทำส่วนที่เครื่องมือเหล่านั้นไม่ทำ: ตัดสิน สืบสวน และลงมือ

ถ้าฉันไม่มีผู้ให้บริการ observability ล่ะ

สัญญาณจากคลาวด์เองครอบคลุมส่วนใหญ่แล้ว: Workers analytics บน Cloudflare, CloudWatch บน AWS และแบบเทียบเท่าในที่อื่น เชื่อมต่อ Datadog, Honeycomb, Axiom, Grafana Cloud หรือ Sentry แล้วชุดข้อมูลเหล่านั้นจะเข้าร่วมกราฟเดียวกันในฐานะแหล่งข้อมูลชั้นหนึ่ง

key queries คืออะไร

คำถามสำหรับมอนิเตอร์ที่แต่ละข้อมี query อยู่เบื้องหลัง สร้างจาก telemetry และโค้ดที่เชื่อมต่อของคุณ และรันกับข้อมูลจริงหนึ่งวันก่อนจะถูกเก็บไว้ คำถามที่บัญชียังตอบไม่ได้จะถูกเก็บไว้อย่างตั้งใจ ในฐานะช่องว่างของ telemetry ที่คุ้มค่าจะปิด

การแก้ instrumentation จะเปลี่ยน stack ของฉันไหม

ไม่ มันเลียนแบบการตั้งค่าที่คุณใช้อยู่แล้วและไม่เพิ่ม dependency ใหม่: event แบบมีโครงสร้าง การจับ error และ attribute ที่การสืบสวนต้องใช้จริง

การมอนิเตอร์ด้วย AI สร้างสัญญาณรบกวนจาก alert เพิ่มไหม

คำตัดสินเริ่มต้นคือไม่มีปัญหา check รู้ว่าทิศทางใดแย่กว่า การเปลี่ยนแปลงในทางที่ดีขึ้นไม่มีทางดัง และทราฟฟิกที่แกว่งอย่างเดียวไม่ถูกนับเป็นความล้มเหลว ความเงียบคือฟีเจอร์: Polylane เอ่ยปากเมื่อมีของจริงพัง

กรณีใช้งานเพิ่มเติม

เลิกจ้องแดชบอร์ด เริ่มอ่านคำตอบ