Get started Dashboard
Monitoring coverage ·
Part 9 of Platform playbooks

How to Monitor a Convex App in Production

Explore with AI

Start with the Health page, the Logs page and the npx convex logs and npx convex insights commands. These only show recent activity, so on a Convex Pro plan you should also stream logs to Axiom, Datadog, PostHog or a webhook, and send exceptions to Sentry or PostHog. Then set alerts on the function_execution events Convex emits for every function run: failure rate per function, slow functions, write conflicts, scheduler lag and functions getting close to their limits.

On this page

To monitor a Convex app in production, start with what every deployment already has: the Health page, the Logs page, and the npx convex logs and npx convex insights commands. All of these show recent activity only. On a Pro plan, add a log stream to Axiom, Datadog, PostHog or a webhook, and turn on exception reporting to Sentry or PostHog. Then build alerts from the function_execution events Convex sends for every function run.

Convex runs your backend as TypeScript functions that sit next to its database. That means almost everything worth watching shows up as a function execution: its status, how long it took, what it read and wrote, and why it ran. This page goes from the built-in views to alerts that reach whoever is on call.

What monitoring a Convex app means

A Convex backend is a set of functions of four types: queries, mutations, actions and HTTP actions. They run on either the default Convex runtime or the Node.js runtime. The scheduler and cron jobs run functions with no user present. Clients subscribe to queries, and Convex reruns those queries whenever the data they read changes.

Convex turns every run into structured events:

  • A function_execution event for each run. It holds the function path and type, status (success or failure), execution_time_ms, an error_message on failure, a usage block (documents and bytes read and written, egress, action memory), and a run_reason.
  • A console event for each console.log, console.warn or console.error, with a log_level and an is_truncated flag.
  • Deployment-level topics: concurrency_stats, scheduler_stats, current_storage_usage, storage_api_bandwidth, audit_log and ai_gateway_usage.

Monitoring Convex in production comes down to three jobs. Keep those events longer than the dashboard does. Write alerts on the few that mean users are hurting. Be able to follow a single failed request from the error a user saw to the log line behind it. I built and led the Workers observability team at Cloudflare, and the rule I’d apply to any hosted backend applies here: get the platform’s own events into the tool you already alert from before you add any custom instrumentation.

What the Convex dashboard and CLI already show you

The Health page

The Health page is where each deployment opens. It shows:

  • Failure rate: the percentage of failed requests, per minute, over the last hour.
  • Cache hit rate: the percentage of cache hits, per minute, over the last hour. This applies to query functions only.
  • Scheduler status: if the scheduler falls behind because too many tasks are scheduled, this card shows the lag in minutes.
  • Last deployed: when functions were last pushed. Check it first when the failure rate jumps.
  • Insights: functions that read too many bytes or too many documents in one transaction, and functions hitting write conflicts. Each insight comes with a chart and an event log.
  • Integrations (Professional only): the status of your log streams and exception reporting.

The Logs page

The Logs page gives a realtime view of function activity plus a short history. Each entry shows the time, request ID, outcome, function name, output and duration in milliseconds. Duration here leaves out network latency. You can filter by text, function, execution status and log severity. Clicking an entry shows every log line that shares its request ID. Because the page doesn’t keep a complete history, older requests may be gone by the time you look.

The CLI

# Tail production logs in the terminal, and keep a copy you can grep later
npx convex logs --prod | tee ./logs.txt

# Health insights for the last 72 hours: write conflicts and resource limit problems
npx convex insights --prod --details

# Check a hunch against production data with a read-only, sandboxed query
npx convex run --prod --inline-query 'await ctx.db.query("messages").take(5)'

# List tables, then print rows from one
npx convex data --prod
npx convex data messages --prod

Inline queries run in a sandbox. They can read data but they can’t modify the database or reach the network, which makes them safe for poking at production during an incident.

The Convex signals worth an alert

The events carry a lot of fields. These are the ones worth paging on or checking every day:

  • Failure rate per function: status equal to failure, divided by total calls, grouped by function path. A single failing mutation matters more than a small rise in the overall rate.
  • Slow functions: execution_time_ms above the level your users will tolerate, grouped by function path.
  • Write conflicts: events carrying occ_info, which names the table, the document and the conflicting function. Also watch mutation_retry_count on successful mutations.
  • Scheduler lag: the scheduler_stats topic, or the scheduler status card on the Health page.
  • Functions nearing limits: console events with a system_code, which Convex adds automatically when a function is approaching its limits.
  • Error-level console logs: log_level equal to ERROR.
  • Usage drift: database_read_documents, database_io_read_bytes and action_memory_used_mb per function, compared week on week.

Step 1: Stream Convex logs to a destination you can alert from

Log streams send function executions, console logs and deployment events to Axiom, Datadog, PostHog or any webhook URL. They need a Convex Pro plan.

  1. Open the production deployment in the Convex dashboard, then go to Settings and the Integrations tab.
  2. Choose a destination and enter what it asks for:
    • Axiom: a dataset name, an API key, and optional attributes that are added to every event.
    • Datadog: your Datadog site location, an API key, and comma-separated tags sent in the ddtags field.
    • PostHog: a project token, an optional host (US Cloud by default, EU Cloud or self-hosted if you set one), and an optional service name. Logs arrive in OpenTelemetry log format.
    • Webhook: a URL. Each request body is a JSON array of events.
  3. Save it and wait for the verification event. Once it arrives, the Convex dashboard shows the Axiom stream as verified and active.
  4. Add attributes or tags that name the project and environment. They make queries across several deployments easier to read.
  5. With Axiom, open the Integrations section of the Dashboards tab. Axiom creates a Convex dashboard for you, covering execution metrics, errors and performance. Fork it if you want to change the layout.

In Axiom, the fields specific to each event sit under a data prefix (data.status, data.function.path). Deployment metadata sits under convex (convex.deployment_name, convex.deployment_type). Webhook payloads carry the same fields at the top level of each event.

Step 2: Send function exceptions to Sentry or PostHog

Logs tell you a function failed. An error tracker groups the failures, keeps the stack traces and tracks which users were affected. Exception reporting is a Pro feature too.

  1. In Sentry, create a project and set its platform to Node.js so Convex exception events are processed correctly. Copy the DSN.
  2. In the Convex dashboard, open Settings, then Integrations, click the Sentry card and paste the DSN. You can add tags of your own if you like.
  3. Convex tags every event with func, func_type, func_runtime, request_id, server_name and environment (prod, dev or preview). When the function is authenticated, it sets the user to the caller’s tokenIdentifier. You can’t override these tags.
  4. Expect exceptions to take a minute or two to show up.

For PostHog Error Tracking, use the PostHog card and your project token. Each $exception event carries convex_function, convex_function_type, convex_deployment_type and convex_request_id, and stack traces are included. For Datadog Error Tracking, set it up through Datadog’s Sentry SDK path, then use the Convex Sentry integration.

Step 3: Turn function_execution events into alerts

Axiom monitors in APL

These queries follow the field names Axiom documents for the Convex integration. The first one pages when a function’s error rate goes above 5%:

['convex']
| where ['data.topic'] == "function_execution"
| where ['convex.deployment_type'] == "prod"
| summarize total_calls = count(), failures = countif(['data.status'] == "failure") by ['data.function.path']
| where total_calls >= 20
| extend error_rate = (failures * 100) / total_calls
| where error_rate > 5

The total_calls floor stops a function called twice in the window from paging you over one failure. Before you rely on the deployment_type filter, check which values your own events carry.

This one finds functions that took longer than 5,000 ms:

['convex']
| where ['data.topic'] == "function_execution" and ['data.execution_time_ms'] > 5000
| summarize count() by ['data.function.path']

This one ranks functions by how much data they read. It is the query to run once a week, and the one to reach for when database I/O costs climb:

['convex']
| where ['data.topic'] == "function_execution"
| summarize avg_docs_read = avg(['data.usage.database_read_documents']), avg_bytes_read = avg(['data.usage.database_io_read_bytes']) by ['data.function.path']
| order by avg_bytes_read desc

Axiom’s own example uses database_read_bytes, but the current Convex event schema calls this field database_io_read_bytes, so the query above uses that name.

Axiom monitors can alert on the average execution time too. For example, you can fire when the average across functions goes above 2 seconds in any 5 minute period, and route the alert to email, Slack, PagerDuty or another notifier.

Datadog and PostHog

In Datadog, write log monitors on the same fields: failed function_execution events grouped by function path, and execution time above your threshold. Use the tags you set in ddtags to keep production separate from everything else. In PostHog, the full event is the log record body, and the deployment metadata comes as resource attributes, so you filter on those.

A webhook receiver of your own

A webhook stream fits when you want the events in your own pipeline or warehouse. Every request is signed. Convex computes an HMAC-SHA256 of the body, encodes it as lowercase hex with a sha256= prefix, and sends it in the x-webhook-signature header. You’ll find the secret in the dashboard when you set up the webhook. Verify the signature in constant time, and reject events whose timestamp is too old:

export default {
  async fetch(req: Request, env: { WEBHOOK_SECRET: string }) {
    const body = await req.arrayBuffer();
    const signature = req.headers.get("x-webhook-signature");
    if (!signature) return new Response("Unauthorized", { status: 401 });

    const key = await crypto.subtle.importKey(
      "raw",
      new TextEncoder().encode(env.WEBHOOK_SECRET),
      { name: "HMAC", hash: "SHA-256" },
      false,
      ["verify"],
    );
    const valid = await crypto.subtle.verify(
      "HMAC",
      key,
      Uint8Array.fromHex(signature.replace("sha256=", "")),
      body,
    );
    if (!valid) return new Response("Unauthorized", { status: 401 });

    const events = JSON.parse(new TextDecoder().decode(body));
    if (events[0].timestamp < Date.now() - 5 * 60 * 1000) {
      return new Response("Request expired", { status: 403 });
    }

    const failures = events.filter(
      (e: any) => e.topic === "function_execution" && e.status === "failure",
    );
    // forward failures to your alerting tool here
    return new Response("ok");
  },
};

Reading write conflicts, retries and query reruns

Some Convex problems never show up as a failed request. Three fields catch them.

occ_info appears when a function hits a write conflict with another function. It names the table_name, the document_id, the write_source (the conflicting function) and a retry_count. When the same document keeps showing up, many callers are writing to one hot document. A counter that every request increments is the usual culprit. The Health page and npx convex insights list the same conflicts.

mutation_retry_count counts failed attempts before a mutation succeeded. The mutation reports success, so an error-rate alert never sees it. Users still wait for every retry. Track the average retry count per mutation next to latency.

run_reason says why a function ran: initialSubscription, dataChange, identityChange, webSocket, httpApi, httpEndpoint, cron, scheduler, action or tester. A query with a high count of dataChange reruns is being invalidated by frequent writes to data it reads. Subscription updates count as function calls in billing, so this field explains a function call bill that grows faster than your traffic. mutation_queue_length does the same job for mutation backlogs within a single client session.

Catching functions before they hit Convex limits

Convex enforces limits that turn a slow-growing problem into a sudden failure. The ones worth knowing:

  • Actions on the Convex runtime get 64 MiB of RAM, and Node.js actions get 512 MiB. Watch action_memory_used_mb for actions creeping toward their runtime’s limit.
  • Function arguments and return values can each be up to 16 MiB. Node actions accept arguments up to 5MiB. The function_args_bytes and function_returns_bytes usage fields show how close you are.
  • A document can be up to 1 MiB, and an array can hold up to 8192 elements. A list you keep appending to inside one document will hit one of these limits eventually.
  • Code size is capped at 32 MiB per deployment.
  • The Free plan includes 1,000,000 function calls a month in total. Professional includes 25,000,000 a month, and calls beyond that are billed.

Convex warns you before some of these fail: it adds a console event with a system_code when a function gets close to a limit. Alert on any console event where system_code is present. The Health page insights flag functions that read too many bytes or documents in one transaction.

Following one failed request from the browser to the log line

In production, Convex hides server-side error details from the browser console so it doesn’t leak server state. What the client gets is a Request ID. Tie your tools together through it:

  1. Log the Request ID in your frontend’s error handler, or pass it to your client-side error tracker.
  2. For a recent failure, paste the ID into the search box on the Logs page. You’ll see every line from that request, including the exception.
  3. For an older failure, search your log stream for function.request_id (under data. in Axiom). In Sentry, search the request_id tag.
  4. With the failing function identified, read the relevant data using npx convex run --prod --inline-query to check whether the documents look the way the code expects.
  5. Compare the failure’s start time with Last deployed on the Health page, and with the deploy history in your CI.

Where Convex monitoring goes wrong

  • Treating the Logs page as your history. It keeps a short window, and the request you need will have gone by the next morning. Fix: set up a log stream before you launch.
  • Streaming in the legacy format. Log streams set up before May 23, 2024 use the legacy event schema, which has different field names. Fix: update the stream to the new format in the dashboard under Settings, Integrations, then update your queries.
  • Letting dev and preview deployments page you. Exception tags and event metadata cover every deployment type. Fix: filter alerts on environment in Sentry and on convex.deployment_type in your log tool.
  • Accepting unsigned webhook traffic. Anyone who learns the URL can post fake events to it. Fix: verify x-webhook-signature in constant time and reject stale timestamps.
  • Sentry events arriving garbled. The Sentry project platform wasn’t set to Node.js. Fix: create the project with the Node.js platform.
  • Error-rate alerts that fire on one failure. A function with two calls and one failure has a 50% error rate. Fix: require a minimum number of calls in the window before the rate counts.
  • Missing slow mutations that succeed. Write conflicts get retried, and the final success hides the cost. Fix: alert on occ_info and on the average mutation_retry_count per mutation.
  • Writing queries against the wrong field paths. Axiom nests event data under data, while webhook payloads don’t. Fix: look at one raw event in your destination before you write a monitor.

Triaging Convex alerts with Polylane

If you connect a Convex deployment as a cloud account in Polylane, its account page has an Alerts tab that triages the account’s alerts automatically. Each alert that fires becomes an issue. A confirmed issue gets a fix run, which traces the cause and, when a code change fixes it, opens a pull request in your repository for you to review.

Running on Convex? See how Polylane monitors Convex in production.

Common questions.

Do I need a paid Convex plan to monitor production?

No for the basics. The Health page, the Logs page, npx convex logs and npx convex insights don't need a log stream or exception integration. Log streams to Axiom, Datadog, PostHog or a webhook, and exception reporting to Sentry or PostHog, both need a Convex Pro plan.

How long does the Convex dashboard keep logs?

The Logs page shows a short history of recent function logs and new ones as they arrive. It isn't a complete historical record, so older requests may be missing. To keep history, set up a log stream, or run npx convex logs --prod | tee ./logs.txt to save a copy locally.

Which log and error destinations does Convex support?

Log streams can go to Axiom, Datadog, PostHog, or a webhook at any URL you choose. Exception reporting supports Sentry and PostHog Error Tracking. Datadog Error Tracking works through its Sentry SDK path combined with the Convex Sentry integration.

How do I find the server error behind a failure a user saw in production?

Convex hides server error details from the browser in production and gives the client a Request ID. Paste that ID into the search box on the Logs page for recent requests. For older ones, search function.request_id in your log stream or the request_id tag in Sentry.

What should my first Convex alert be?

Start with a failure-rate alert per function: count function_execution events with status failure, divide by total calls, group by function path, and fire above 5%. Set a minimum call count so low-traffic functions can't page you over a single error. Add a slow-function alert on execution_time_ms next.

How do I monitor write conflicts in Convex?

Failed or retried functions carry an occ_info object with the table name, document ID and the conflicting function. Alert on events where it's present, and track mutation_retry_count on successful mutations. Run npx convex insights --prod --details to see conflicts from the last 72 hours.

Why does my Convex log stream use different field names from the docs?

Log streams configured before May 23, 2024 send the legacy event schema. Convex recommends updating to the current format. Once you do, update your queries and monitors to the new field names.

How quickly do Convex exceptions appear in Sentry or PostHog?

Convex documents a delay of a minute or two for both Sentry and PostHog. For anything faster, alert from the log stream on failed function_execution events, which carry the error message and stack trace.

Sources

  1. Log Streams, Convex Developer Hub
  2. Legacy Event Schema, Convex Developer Hub
  3. Exception Reporting, Convex Developer Hub
  4. Health, Convex Developer Hub
  5. Logs, Convex Developer Hub
  6. CLI, Convex Developer Hub
  7. Limits, Convex Developer Hub
  8. Observing your app in production, Stack by Convex
  9. Connect Axiom with Convex, Axiom Docs
  10. Send data from Convex to Axiom, Axiom Docs
  11. Convex + Axiom: Complete observability for reactive backends
  12. Polylane documentation

About the author

Boris Tane

Founder of Polylane

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

More in this series

Platform playbooks

  1. 1 How to Monitor a Django App on Render
  2. 2 How to Monitor a FastAPI App on Railway: Logs, Traces and Alerts
  3. 3 How to Monitor a Supabase App in Production
  4. 4 How to Debug Cloudflare Workers Errors: Logs, Traces and Error Codes
  5. 5 How to debug Vercel function timeouts
  6. 6 Vercel 504 Gateway Timeout on Serverless Functions: Causes and Fixes
  7. 7 Cloudflare Workers error 1101: causes and how to fix it
  8. 8 Cloudflare Hyperdrive connection errors: causes and fixes
  9. 9 How to Monitor a Convex App in Production

Related

Nobody should be on-call. Polylane watches your infra, finds what broke, and writes the fix.

Get started for free