แดชบอร์ด

งานโครงสร้างพื้นฐานที่ไม่มีใครถูกจ้างมาทำ วาดแผนที่ เฝ้าดู และรักษาความถูกต้องโดยเอเจนต์

ลงชื่อเข้า waitlist

กราฟเดียว ทุกผู้ให้บริการ ไม่ใช่แผนภาพที่คุณต้องดูแล: กราฟที่ซิงก์ตัวเอง

Topology
Search your cloud resources...
Galaxy
checkout-edge Worker
Critical Critical to your architecture.

Issue hotspot: 4 issues in the last 7 days

Change hotspot: 12 changes in the last 7 days

Cloudflare · coreplane · earth · Compute
hd-prod Hyperdrive
Critical Critical to your architecture.

Change hotspot: 7 changes in the last 7 days

Cloudflare · coreplane · earth · Databases
ingest-events Lambda function
Standard Important to your architecture.

Issue hotspot: 2 issues in the last 7 days

AWS · coreplane-prod · us-east-1 · Compute
payments-db Database
Critical Critical to your architecture.

Issue hotspot: 1 issue in the last 7 days

Change hotspot: 3 changes in the last 7 days

PlanetScale · coreplane · us-east · Databases

ทุกการเปลี่ยนแปลงมีบันทึก สิบสองนาทีก่อน regression มีคนแตะคิว

Clouds prod-aws Changes

SQS visibility timeout lowered on checkout-events

Moderate impact configuration ·Synced 12 minutes ago

VisibilityTimeout on checkout-events dropped from 120s to 15s. Consumers that hold a message longer than 15 seconds will see it delivered twice; the dead-letter queue threshold is unchanged.

What we're watching
ApproximateAgeOfOldestMessage
checkout-events · baseline at change 3.2s · worse if up
Watching: no issue since this change
NumberOfMessagesReceived
checkout-events · baseline at change 41/min · worse if up
Watching: no issue since this change
Triggering events
SetQueueAttributes CloudTrail · deploy-bot · 12:41:02Z attached to this record's delta window
Full diff (3 changes)
Nodes (2) Edges (1)
2 nodes modified · 1 edge removed

ส่วนที่เงียบ ถูกเฝ้าดู cron ที่ค้าง backlog ที่โตขึ้น worker ที่หยุดไป

Issues reconcile-orders stopped running Overview

reconcile-orders stopped running

High Incident ·Detected by Polylane
OverviewMetricsLogsTracesTimelineProperties
What changed
  • runs/day 4 → 0: no successful run since Tue 02:00 UTC
  • last attempt exited with code 137 after 91s (Tue 02:01)
  • peak memory on the final run was 2.4× the previous seven-day average
Why it matters

Orders placed since Tuesday have no reconciliation record. The finance export and the daily revenue report both read from the table this job writes.

Suggested investigation steps
  1. 1. Read the logs from the last attempt on reconcile-orders
  2. 2. Check for memory limit or plan changes on the service in the change records
  3. 3. Confirm the schedule still exists in render.yaml on main
Investigate Detected 02:31 UTC · 30 minutes after the missed run

Polylane ทำงานอย่างไร มันเรียนรู้ระบบของคุณ เฝ้าดู สืบสวน และลงมือ

  • มันเรียนรู้ระบบของคุณก่อน

    context graph วาดแผนที่ทุกทรัพยากรและการพึ่งพาทั่วคลาวด์ repo และผู้ให้บริการ observability ของคุณ เอเจนต์ใช้เหตุผลบน topology จริง ไม่ใช่การเดา

  • การตรวจจับโดยไม่ใช้ threshold

    check ในตัวสำหรับทุกผู้ให้บริการ บวก check ที่สร้างจาก query และแดชบอร์ดที่คุณบันทึกไว้เอง การประมวลผลทางสถิติและเอเจนต์ตัดสินร่วมกัน และการเปลี่ยนแปลงในทางที่ดีขึ้นไม่มีทางเปิด issue

  • การสืบสวนที่แสดงหลักฐาน

    ทุกข้ออ้างลิงก์กลับไปยัง query, บรรทัด log หรือบันทึกการเปลี่ยนแปลงที่อยู่เบื้องหลัง คำตัดสินที่ไม่มีหลักฐานจะถอยกลับเป็นสรุปไม่ได้

  • การเขียนต้องได้รับสิทธิ์ ไม่ใช่ถือว่ามี

    บัญชีเชื่อมต่อแบบอ่านอย่างเดียว rollback ปิดอยู่โดยค่าเริ่มต้น จำกัดอัตรา และมีบันทึกไว้ การเปลี่ยนโค้ดผ่านการรีวิวตามปกติของคุณ

  • มันคมขึ้นทุกสัปดาห์

    memory โน้ตประจำวัน และ query สำหรับมอนิเตอร์ที่ยืนยันซ้ำกับข้อมูลจริง: การสืบสวนเดือนกรกฎาคมเรียนรู้จากเดือนมิถุนายน

มันเสียบเข้ากับสิ่งที่คุณรันอยู่แล้ว เชื่อมต่อแบบอ่านอย่างเดียวแล้วเริ่มได้เลย

คำถาม

วันนี้ Polylane รองรับผู้ให้บริการใดบ้าง

AWS, Cloudflare, Vercel, Render, Fly.io, Kubernetes, PlanetScale, Supabase และ Modal บวก GitHub สำหรับโค้ด และ Datadog, Honeycomb, Axiom, Grafana Cloud และ Sentry สำหรับ telemetry หน้าการผสานรวมติดตามแคตตาล็อกฉบับเต็ม รวมถึงสิ่งที่กำลังจะมา

มันเชื่อมโยงทรัพยากรข้ามคลาวด์อย่างไร

ตามที่ทราฟฟิกไหลจริง: hostname, DNS record และ IP address ที่จับคู่ข้ามผู้ให้บริการ ตัวแปร environment ที่เผยการพึ่งพา trace ในที่ที่คุณมี และ infrastructure-as-code ที่ประกาศว่าอะไร deploy ไปที่ไหน

ฉันต้องติดแท็กทรัพยากรหรือวาด topology เองไหม

ไม่ เชื่อมต่อแต่ละบัญชีแบบอ่านอย่างเดียว แล้วกราฟจะสร้างและดูแลตัวเอง คุณแก้ไขหรือใส่คำอธิบายอะไรก็ได้ และเอเจนต์คอยอัปเดตให้เป็นปัจจุบันทุกการซิงก์

แต่ละคลาวด์ต้องให้สิทธิ์เข้าถึงมากแค่ไหน

ขั้นต่ำ อ่านอย่างเดียวโดยค่าเริ่มต้น: AWS ผ่าน role ของ CloudFormation ที่จำกัดขอบเขต Cloudflare ผ่านสิทธิ์โทเค็นที่กรอกไว้ล่วงหน้า และแบบเทียบเท่าในที่อื่น ๆ สิทธิ์เขียนเป็นการตัดสินใจแยกต่างหาก เป็นรายบัญชี

มันเฝ้าดูระบบเบื้องหลังใดบ้าง

อะไรก็ตามที่ผู้ให้บริการของคุณรัน: คิว SQS และ job ตามกำหนดเวลาบน AWS, Cloudflare Queues, background worker และ cron job ของ Render, CronJob ของ Kubernetes, เครื่องของ Fly.io ถ้ามันอยู่ในบัญชีที่เชื่อมต่อไว้ มันก็อยู่ในกราฟ

กรณีใช้งานเพิ่มเติม

เชื่อมต่อคลาวด์ของคุณ กราฟสร้างตัวเอง