Get started Dashboard
Monitoring coverage · Updated
Part 3 of Platform playbooks

How to Monitor a Supabase App in Production

Explore with AI

Scrape the project's Metrics API into Prometheus or another compatible tool, send its logs to your alerting tool with a log drain, run Supabase's SQL detection checks on a schedule, and add trace headers to your app's requests so its traces line up with Supabase's logs. Then alert on a short list of signals: API and Auth server errors, connection pressure, blocked or long-running queries, slow statements, replication lag and database growth.

On this page

To monitor a Supabase app in production, get four kinds of signal out of the project and into the tool you already alert from. Those are database metrics from the Metrics API, platform logs through a log drain, scheduled SQL checks against Postgres, and traces from your own app that carry a trace ID into Supabase’s logs. Then write a short list of alerts on the signals that actually break apps: server errors on the API and Auth, connection pressure, blocked or long-running queries, slow statements, replication lag and database growth.

The dashboards and Logs Explorer in Supabase Studio work well for a quick health check. They live apart from your app’s own telemetry, though, so you can’t line up a latency spike in your API with a saturated connection pool. The steps below take a project from dashboards you have to go and look at to alerts that come to you.

What monitoring a Supabase app covers

A Supabase project is several services behind one URL: Postgres, the API gateway and PostgREST, Auth, Storage, Edge Functions and Realtime. Your app adds a frontend and usually a server of its own. Each part fails in its own way, and Supabase exposes each one through a different route:

  • The Metrics API publishes roughly 200 Postgres performance and health metrics in Prometheus exposition format. They cover CPU and memory, disk I/O and WAL, Supavisor and Postgres connection pools, and query performance. The whole set refreshes every minute.
  • Logs come from every subsystem and can be queried with SQL in Explorer. Log drains stream them to another tool on the Pro, Team and Enterprise plans.
  • Live Postgres statistics come from pg_stat_activity and pg_stat_statements. You read them with the Supabase CLI, SQL in Explorer, or execute_sql on the Supabase MCP server.
  • Advisors return security and performance findings, each with a severity, the affected object and a documentation link.
  • Client-side tracing in the Supabase SDKs adds W3C trace headers to outgoing requests, so a trace_id from your app appears in API Gateway and Edge Function logs.

One detail shapes everything below. pg_stat_activity is a live snapshot. pg_stat_statements and the cache counters are cumulative since their last reset. A snapshot tells you what is happening right now. A cumulative counter only tells you about a time window once you compare two saved readings.

I built and led the Workers observability team at Cloudflare, and the advice I would give for any hosted platform applies here too: get its telemetry into the tool that already holds your app’s traces and alerts before you write a single alert rule.

The signals that break Supabase apps first

With about 200 metrics on offer, most teams start with too many charts and no pages at all. These are the signals worth an alert or a scheduled check:

  • Server errors on the API and Auth. The 5xx rate on API Gateway and Auth, each measured on its own. This is the closest thing to what users feel.
  • Access failures. A jump in 401 and 403 responses usually points at a broken client flow or a policy change.
  • Connection pressure. Client connections against max_connections. Reserved slots, role limits and pooler limits can cut a client off before the instance-wide number looks full.
  • Blocked and long-running sessions. Queries active for a long time, transactions left idle, and sessions waiting on another session’s lock.
  • Slow statements. Mean execution time per statement from pg_stat_statements, compared hour on hour. Deploys that drop an index show up here first.
  • CPU, memory and disk I/O. Supabase suggests a right-sizing alert when CPU or memory stays above 80%. High I/O wait often means missing indexes or queries that scan too much data.
  • Replication lag, if you run read replicas.
  • Database size growth, which warns you well before the disk fills.
  • Advisor findings at WARN or ERROR severity.

Step 1: Scrape the Metrics API into Prometheus

Every hosted project has a metrics endpoint at https://<project-ref>.supabase.co/customer/v1/privileged/metrics, protected with HTTP Basic Auth and a secret API key. If you use Grafana Cloud, the integration in the Supabase Dashboard sets this up for you. Datadog can read the same endpoint through the Datadog Agent’s OpenMetrics integration, and AWS Managed Prometheus works as well. For a self-hosted Prometheus and Grafana:

  1. Copy a secret API key (it starts with sb_secret_) into your secret manager or an environment variable.

  2. Check that the endpoint answers:

    curl -s -u "username:$SUPABASE_SECRET_KEY" \
      "https://$PROJECT_REF.supabase.co/customer/v1/privileged/metrics" | head -n 20

    You should see plain-text lines in Prometheus exposition format, one metric per line. An authentication error means the key or the username is wrong.

  3. Add a scrape job to prometheus.yml:

    scrape_configs:
      - job_name: 'supabase'
        scrape_interval: 60s
        metrics_path: /customer/v1/privileged/metrics
        scheme: https
        basic_auth:
          username: username
          password: '<secret API key (sb_secret_...)>'
        static_configs:
          - targets:
              - '<project-ref>.supabase.co:443'
            labels:
              project: '<project-ref>'
  4. Keep scrape_interval at 60 seconds to match the refresh. Add one job per project ref so each project’s labels stay separate. If Prometheus sits behind a proxy, allow outbound HTTPS to *.supabase.co.

  5. In Grafana, add Prometheus as a data source under Connections, Data sources. Then import dashboard.json from the supabase-grafana repository under Dashboards, New, Import. That gives you over 200 panels covering CPU, I/O, WAL, replication, index bloat and query throughput.

Step 2: Turn the metrics into alerts that page someone

The supabase-grafana repository ships example PromQL rules. Three of them are worth loading on day one:

groups:
  - name: supabase
    rules:
      - alert: PostgresDatabaseDown
        expr: pg_up == 0
        for: 5m
        labels:
          severity: critical
      - alert: PostgresReplicationLagHigh
        expr: (physical_replication_lag_physical_replication_lag_seconds > 600) and (rate(physical_replication_lag_physical_replication_lag_seconds[10m]) > 0)
        for: 5m
        labels:
          severity: warning
      - alert: PostgresDatabaseSizeGrowth
        expr: (pg_database_size_mb - pg_database_size_mb offset 12h) / pg_database_size_mb offset 12h * 100 > 20
        for: 5m
        labels:
          severity: warning

The first fires when the instance is still sending metrics but the database has been down for 5 minutes. The second fires when replication lag passes 10 minutes and is still climbing. The third fires when the database has grown by 20% or more in 12 hours.

On top of those, add a connection saturation alert that fires as pools near capacity, so you have a few minutes to investigate before connections start failing. Add a right-sizing alert for CPU or memory held above 80%, which gives you time to upgrade before users notice. Set thresholds for disk use and long-running transactions to fit your project’s size. Route everything through Alertmanager, Grafana OnCall, PagerDuty or whichever tool already wakes your on-call.

Starter alert thresholds for a Supabase project
SignalStarting thresholdSource of the default
Database down pg_up equals 0 for 5 minutes supabase-grafana example alerts
Replication lag Over 600 seconds and still rising supabase-grafana example alerts
Database size 20% or more growth in 12 hours supabase-grafana example alerts
CPU or memory Held above 80% Supabase Metrics API guidance
Client connections 80% of max_connections Supabase detection checks
API or Auth server errors 20 or more errors, at least 1%, and double the previous hour Supabase detection checks
Slow statement Mean of 100 ms or more and double the previous hour, with 20 or more calls Supabase detection checks
Session age Active or idle in transaction for over 30 seconds Supabase detection checks

Step 3: Drain Supabase logs into the tool you alert from

Metrics tell you something is wrong. Logs tell you which request failed. Log drains send the logs of the whole stack (Postgres, Auth, Storage, Edge Functions and the rest) to one or more destinations. Set them up under Project Settings, Log Drains. They are available on the Pro, Team and Enterprise plans, and Supabase bills them per drain, per million events and for egress.

The destinations are a custom HTTP endpoint, OpenTelemetry over OTLP, Datadog, Grafana Loki, Amazon S3, Sentry, Axiom, Last9 and Syslog. HTTP destinations receive batched POST requests of up to 250 events, or whatever has arrived within 1 second, whichever comes first. A few details matter when you pick one:

  • OTLP needs an endpoint that accepts application/x-protobuf at /v1/logs. It works with the OpenTelemetry Collector, Grafana Cloud, New Relic, Honeycomb, Datadog’s OTLP ingestion and Elastic.
  • Loki must accept structured metadata, and you should raise the maximum number of structured metadata fields to at least 500 so large events are not rejected.
  • Sentry receives the events in its Logging product. They do not arrive as Sentry errors.
  • Custom endpoints currently receive unsigned requests, so add a secret header and check it on your side.
  • S3 suits long-term retention and compliance archives next to a live destination.

Once logs arrive, build alerts in that tool on error-level events from Auth, Edge Functions and Postgres. A drop in volume is worth an alert too.

Step 4: Schedule Supabase’s SQL detection checks

Supabase publishes a set of detection checks for health, security, performance and capacity. Its own monitoring agents use the same checks. Each check returns finding, clear, or unable to assess along with the input that was missing. Run them from a cron job, a scheduled CI workflow or an agent with project-scoped, read-only access (the Supabase MCP server takes read_only=true). Record the observation time, the window and the thresholds with every result.

Server errors on the API and Auth

Run this against the logs source for the last complete UTC hour, then again for the hour before it:

select source,
  countIf(status between 100 and 599) as responses,
  countIf(status between 500 and 599) as server_errors,
  countIf(status in (401, 403)) as access_failures
from (
  select source,
    toInt32OrNull(if(source = 'edge_logs',
      log_attributes['response.status_code'], log_attributes['status'])) as status
  from logs
  where source in ('edge_logs', 'auth_logs')
)
group by source;

Work out the error rate per source. Supabase’s starting policy reports a finding at 20 or more server errors, a rate of at least 1%, and at least double the previous hour’s rate. Both windows need at least 100 responses, or the result is unable to assess. Run the same query over the last complete day to track 401 and 403 access failures against the day before.

Connection pressure

select
  count(*) filter (where backend_type = 'client backend') as client_connections,
  current_setting('max_connections')::int as max_connections
from pg_stat_activity;

Report a finding when client connections reach 80% of max_connections. If connections crowd the limit often, move chatty or bursty workloads to the dedicated pooler with sensible pool sizes.

Blocked and long-running sessions

select pid, usename as role, state,
  now() - query_start as query_age,
  now() - xact_start as transaction_age,
  pg_blocking_pids(pid) as blocking_pids
from pg_stat_activity
where datname = current_database()
  and pid <> pg_backend_pid()
  and (
    (state = 'active' and now() - query_start > interval '30 seconds')
    or (state like 'idle in transaction%' and now() - xact_start > interval '30 seconds')
    or cardinality(pg_blocking_pids(pid)) > 0
  )
order by query_start
limit 20;

Every row needs a look. A non-empty blocking_pids names the blockers. If you get 20 rows back, the list may be cut short.

Slow statements

With pg_stat_statements enabled, save a snapshot of calls and total_exec_time per (dbid, userid, queryid, toplevel) every hour. Divide the change in total time by the change in calls for each interval. Supabase’s policy flags a statement whose mean reaches 100 ms and doubles the previous hour, with at least 20 calls in each interval. Throw out comparisons across a reset, an upgrade or an eviction. When a statement is flagged, read its query plan.

Advisors

Call get_advisors with type: "security" and again with type: "performance". Report WARN and ERROR findings with the lint name, the affected object and the documentation link. Keep INFO as context. Run the advisor again after each fix. An empty result does not prove the project is secure.

From the terminal

The Supabase CLI wraps many of the same statistics:

supabase link --project-ref <project-id>
supabase inspect db long-running-queries
supabase inspect db blocking
supabase inspect db outliers
supabase inspect db role-connections
supabase inspect db replication-slots

For disk pressure, bloat, vacuum-stats, table-sizes and index-sizes show where the space went. For query work, unused-indexes, index-usage and seq-scans point at indexing problems.

Step 5: Carry your app’s trace IDs into Supabase logs

Metrics and logs cover the Supabase side. To follow one slow page load from the browser through PostgREST or an Edge Function, the request needs a trace ID that both sides record. The Supabase JS, Swift, Dart and Python SDKs can attach W3C traceparent, tracestate and baggage headers, and the resulting trace_id appears in API Gateway and Edge Function logs.

In JavaScript you need @supabase/supabase-js 2.106.0 or later, @opentelemetry/api at runtime, and a registered tracer provider. From 2.112.0, load the tracing runtime once at your entry point:

import '@supabase/supabase-js/tracing'

import { trace } from '@opentelemetry/api'
import { createClient } from '@supabase/supabase-js'

const supabase = createClient(SUPABASE_URL, SUPABASE_KEY, {
  tracePropagation: true,
})

const tracer = trace.getTracer('my-app')

await tracer.startActiveSpan('load-orders', async (span) => {
  const { data, error } = await supabase.from('orders').select('*')
  span.end()
})

The SDK only adds trace headers to requests for *.supabase.co, *.supabase.in and localhost. When you call Edge Functions from the browser, add traceparent, tracestate and baggage to the function’s CORS allow-list, or import corsHeaders from npm:@supabase/supabase-js@^2.112.3/cors, then redeploy.

Vendor SDKs differ. The OpenTelemetry SDK, and Honeycomb, Grafana or New Relic over OTLP, work with no change. Supabase’s tracing guide lists propagateTraceparent: true as a requirement for Sentry. In the browser it also needs your project URL in tracePropagationTargets, plus sentry-trace in the Edge Function’s CORS allow-list. Datadog’s dd-trace for Node.js adds W3C headers by itself. Datadog Browser RUM needs your project URL in allowedTracingUrls with the tracecontext propagator.

Where Supabase monitoring setups go wrong

Scraping faster than the data changes

The Metrics API refreshes once a minute. A shorter scrape interval stores repeated values and costs storage for nothing. Fix: keep scrape_interval: 60s.

Reading silence as health

If the scrape job breaks, pg_up == 0 never fires, because no data arrives at all. The same goes for logs: Supabase’s checks say zero recorded events alone does not prove service health, and a missing source row calls for a capture check. Fix: alert when the scrape job fails and when a log source goes quiet. Treat an unable-to-assess result as something to act on.

Adding API and Auth errors together

API Gateway events and Auth events are separate observations of different requests. Summing them blurs both rates. Fix: compute and alert on each source by itself, and report unknown status codes separately.

Resetting statistics to get a clean baseline

Resetting pg_stat_statements wipes the history every comparison depends on. Fix: save hourly snapshots and compare deltas. Start collecting now if you have no history.

Cancelling a query because it is old

A long query or a wait event alone does not make a session a blocker, and query age is a different thing from lock-wait time. Fix: act on non-empty blocking_pids, find out what the transaction is for, and run the check again to confirm the problem has cleared.

Trace IDs that never appear in the logs

The SDK never throws when it can’t propagate, so a bad setup fails quietly. The usual causes are a missing @supabase/supabase-js/tracing import on 2.112.0 or later, a call made outside any active span, no registered tracer provider, a Sentry setup without propagateTraceparent: true, or the CDN build, which cannot load the tracing runtime. Fix: look for the SDK’s one-time console warning, wrap calls in spans, and check each item on that list in turn.

Leaving the secret key in a config file

The Metrics API authenticates with a secret (service-role) API key. Fix: inject it from a secret manager or environment variable, rotate it on a schedule, and update Prometheus when you do.

Building monitoring tables inside the project

Writing snapshots back into the database you are watching adds load and tangles the two together. Fix: Supabase’s guidance is to keep saved readings in persistent storage outside the project, or in a historical metrics source you are allowed to use.

Planning on log drains from the Free plan

Drains need the Pro, Team or Enterprise plan. Fix: until you upgrade, rely on the Metrics API and scheduled log queries run through Explorer or the MCP query_logs tool.

A checklist before the app takes real traffic

  1. Metrics API scraped every 60 seconds, one job per project, with the key in a secret manager.
  2. Supabase Grafana dashboard imported, or the Grafana Cloud or Datadog equivalent set up.
  3. Alerts live for database down, replication lag, size growth, connection saturation and sustained CPU or memory, routed to on-call.
  4. An alert on the scrape job failing and on log sources going quiet.
  5. Log drain sending to your alerting tool, with S3 added if you need an archive.
  6. Hourly API and Auth error checks, a daily access-failure check, and connection and blocker snapshots, all running read-only.
  7. pg_stat_statements enabled, with hourly snapshots saved outside the project.
  8. Security and performance advisors run on a schedule, with WARN and ERROR findings sent to a person.
  9. Trace propagation on in the app, with trace headers in the Edge Function CORS allow-lists.

Letting Polylane watch the Supabase project

Polylane connects a Supabase organisation through OAuth or a personal access token and syncs every project’s databases, edge functions, branches, buckets, auth and storage every 15 minutes, with logs and checks on each. It links the auth, storage, realtime and REST services to the Postgres database they sit in front of, triages the account’s alerts, and reruns a set of key questions on a schedule, each paired with a query in the provider’s own language. Accounts start read-only, and when a confirmed cause is a code defect the fix arrives as a pull request for you to review.

Running on Supabase? See how Polylane monitors Supabase in production.

Common questions.

Does Supabase have built-in monitoring?

Yes. Studio has observability dashboards, a Logs Explorer you can query with SQL, reports and Advisors, and the CLI reads live Postgres statistics with `supabase inspect db`. For alerting and for lining up database signals with your app's telemetry, export the data through the Metrics API and log drains.

What is the Supabase Metrics API endpoint and how often should I scrape it?

Each hosted project exposes `https://<project-ref>.supabase.co/customer/v1/privileged/metrics`, protected with HTTP Basic Auth and a secret API key. It publishes roughly 200 Postgres metrics in Prometheus format and refreshes every minute, so scrape it every 60 seconds.

Which Supabase plans include log drains?

Log drains are available on the Pro, Team and Enterprise plans, under Project Settings, Log Drains. Supabase bills them per drain, per million events and for egress. On the Free plan, use the Metrics API and scheduled log queries.

Can I monitor Supabase with Datadog or Grafana Cloud?

Yes. Grafana Cloud has an integration in the Supabase Dashboard. Datadog can read the Metrics API through the Datadog Agent's OpenMetrics integration, and it is also a log drain destination that needs an API key and a region.

What should I alert on first?

Start with the database being down for 5 minutes, replication lag over 10 minutes and rising, and database size growing 20% in 12 hours, all from the supabase-grafana example rules. Then add client connections at 80% of max_connections, CPU or memory held above 80%, and API or Auth server errors of 20 or more at a rate of at least 1% that doubles the previous hour.

How do I find which query is slowing the database down?

Enable `pg_stat_statements`, save hourly snapshots, and compare the mean execution time per statement. Supabase's check flags a mean of 100 ms or more that doubles the previous hour with at least 20 calls. For a quick look, run `supabase inspect db outliers` and `supabase inspect db long-running-queries`, then read the query plan.

How do I connect my frontend traces to Supabase logs?

Use `@supabase/supabase-js` 2.106.0 or later, import `@supabase/supabase-js/tracing` at your entry point from 2.112.0, and create the client with `tracePropagation: true` while an OpenTelemetry span is active. The trace ID then appears in API Gateway and Edge Function logs. Add the trace headers to your Edge Functions' CORS allow-list.

Is it safe to let an AI agent run these monitoring checks?

Supabase's detection checks are written for read-only access: use the project-scoped Supabase MCP server with `read_only=true`, which runs all SQL as a read-only Postgres user. Also restrict the server's feature groups to the ones the checks need. Keep any fix behind human review.

Sources

  1. Own Your Observability: Supabase Metrics API
  2. Supabase Docs: Detection checks
  3. Supabase Docs: Inspect the database
  4. Supabase Docs: Metrics API with Prometheus and Grafana (self-hosted)
  5. Supabase Docs: Log Drains
  6. Supabase Docs: Client-side tracing
  7. Supabase Docs: Model context protocol (MCP)
  8. supabase-grafana: Example alerts
  9. Polylane documentation
  10. Polylane full content

About the author

Boris Tane

Founder of Polylane

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

More in this series

Platform playbooks

  1. 1 How to Monitor a Django App on Render
  2. 2 How to Monitor a FastAPI App on Railway: Logs, Traces and Alerts
  3. 3 How to Monitor a Supabase App in Production
  4. 4 How to Debug Cloudflare Workers Errors: Logs, Traces and Error Codes
  5. 5 How to debug Vercel function timeouts
  6. 6 Vercel 504 Gateway Timeout on Serverless Functions: Causes and Fixes
  7. 7 Cloudflare Workers error 1101: causes and how to fix it
  8. 8 Cloudflare Hyperdrive connection errors: causes and fixes
  9. 9 How to Monitor a Convex App in Production

Related

Nobody should be on-call. Polylane watches your infra, finds what broke, and writes the fix.

Get started for free