# How to Monitor a Convex App in Production

> Monitor a Convex app in production: use the Health page and CLI, stream logs to Axiom or Datadog, report exceptions to Sentry, and alert on failures.

Part 9 of [Platform playbooks](https://polylane.com/series/platform-playbooks/)
By Boris Tane, Founder of Polylane · Published September 30, 2026 · 13 min read
Canonical: https://polylane.com/learn/monitoring-coverage/how-to-monitor-a-convex-app-in-production/

Start with the Health page, the Logs page and the npx convex logs and npx convex insights commands. These only show recent activity, so on a Convex Pro plan you should also stream logs to Axiom, Datadog, PostHog or a webhook, and send exceptions to Sentry or PostHog. Then set alerts on the function_execution events Convex emits for every function run: failure rate per function, slow functions, write conflicts, scheduler lag and functions getting close to their limits.

To monitor a Convex app in production, start with what every deployment already has: the Health page, the Logs page, and the `npx convex logs` and `npx convex insights` commands. All of these show recent activity only. On a Pro plan, add a log stream to Axiom, Datadog, PostHog or a webhook, and turn on exception reporting to Sentry or PostHog. Then build alerts from the `function_execution` events Convex sends for every function run.

Convex runs your backend as TypeScript functions that sit next to its database. That means almost everything worth watching shows up as a function execution: its status, how long it took, what it read and wrote, and why it ran. This page goes from the built-in views to alerts that reach whoever is on call.

## What monitoring a Convex app means

A Convex backend is a set of functions of four types: queries, mutations, actions and HTTP actions. They run on either the default Convex runtime or the Node.js runtime. The scheduler and cron jobs run functions with no user present. Clients subscribe to queries, and Convex reruns those queries whenever the data they read changes.

Convex turns every run into structured events:

- A `function_execution` event for each run. It holds the function path and type, `status` (`success` or `failure`), `execution_time_ms`, an `error_message` on failure, a `usage` block (documents and bytes read and written, egress, action memory), and a `run_reason`.
- A `console` event for each `console.log`, `console.warn` or `console.error`, with a `log_level` and an `is_truncated` flag.
- Deployment-level topics: `concurrency_stats`, `scheduler_stats`, `current_storage_usage`, `storage_api_bandwidth`, `audit_log` and `ai_gateway_usage`.

Monitoring Convex in production comes down to three jobs. Keep those events longer than the dashboard does. Write alerts on the few that mean users are hurting. Be able to follow a single failed request from the error a user saw to the log line behind it. I built and led the Workers observability team at Cloudflare, and the rule I'd apply to any hosted backend applies here: get the platform's own events into the tool you already alert from before you add any custom instrumentation.

## What the Convex dashboard and CLI already show you

### The Health page

The Health page is where each deployment opens. It shows:

- **Failure rate**: the percentage of failed requests, per minute, over the last hour.
- **Cache hit rate**: the percentage of cache hits, per minute, over the last hour. This applies to query functions only.
- **Scheduler status**: if the scheduler falls behind because too many tasks are scheduled, this card shows the lag in minutes.
- **Last deployed**: when functions were last pushed. Check it first when the failure rate jumps.
- **Insights**: functions that read too many bytes or too many documents in one transaction, and functions hitting write conflicts. Each insight comes with a chart and an event log.
- **Integrations** (Professional only): the status of your log streams and exception reporting.

### The Logs page

The Logs page gives a realtime view of function activity plus a short history. Each entry shows the time, request ID, outcome, function name, output and duration in milliseconds. Duration here leaves out network latency. You can filter by text, function, execution status and log severity. Clicking an entry shows every log line that shares its request ID. Because the page doesn't keep a complete history, older requests may be gone by the time you look.

### The CLI

```bash
# Tail production logs in the terminal, and keep a copy you can grep later
npx convex logs --prod | tee ./logs.txt

# Health insights for the last 72 hours: write conflicts and resource limit problems
npx convex insights --prod --details

# Check a hunch against production data with a read-only, sandboxed query
npx convex run --prod --inline-query 'await ctx.db.query("messages").take(5)'

# List tables, then print rows from one
npx convex data --prod
npx convex data messages --prod
```

Inline queries run in a sandbox. They can read data but they can't modify the database or reach the network, which makes them safe for poking at production during an incident.

## The Convex signals worth an alert

The events carry a lot of fields. These are the ones worth paging on or checking every day:

- **Failure rate per function**: `status` equal to `failure`, divided by total calls, grouped by function path. A single failing mutation matters more than a small rise in the overall rate.
- **Slow functions**: `execution_time_ms` above the level your users will tolerate, grouped by function path.
- **Write conflicts**: events carrying `occ_info`, which names the table, the document and the conflicting function. Also watch `mutation_retry_count` on successful mutations.
- **Scheduler lag**: the `scheduler_stats` topic, or the scheduler status card on the Health page.
- **Functions nearing limits**: `console` events with a `system_code`, which Convex adds automatically when a function is approaching its limits.
- **Error-level console logs**: `log_level` equal to `ERROR`.
- **Usage drift**: `database_read_documents`, `database_io_read_bytes` and `action_memory_used_mb` per function, compared week on week.

## Step 1: Stream Convex logs to a destination you can alert from

Log streams send function executions, console logs and deployment events to Axiom, Datadog, PostHog or any webhook URL. They need a Convex Pro plan.

1. Open the production deployment in the Convex dashboard, then go to **Settings** and the **Integrations** tab.
2. Choose a destination and enter what it asks for:
   - **Axiom**: a dataset name, an API key, and optional attributes that are added to every event.
   - **Datadog**: your Datadog site location, an API key, and comma-separated tags sent in the `ddtags` field.
   - **PostHog**: a project token, an optional host (US Cloud by default, EU Cloud or self-hosted if you set one), and an optional service name. Logs arrive in OpenTelemetry log format.
   - **Webhook**: a URL. Each request body is a JSON array of events.
3. Save it and wait for the `verification` event. Once it arrives, the Convex dashboard shows the Axiom stream as verified and active.
4. Add attributes or tags that name the project and environment. They make queries across several deployments easier to read.
5. With Axiom, open the **Integrations** section of the **Dashboards** tab. Axiom creates a Convex dashboard for you, covering execution metrics, errors and performance. Fork it if you want to change the layout.

In Axiom, the fields specific to each event sit under a `data` prefix (`data.status`, `data.function.path`). Deployment metadata sits under `convex` (`convex.deployment_name`, `convex.deployment_type`). Webhook payloads carry the same fields at the top level of each event.

## Step 2: Send function exceptions to Sentry or PostHog

Logs tell you a function failed. An error tracker groups the failures, keeps the stack traces and tracks which users were affected. Exception reporting is a Pro feature too.

1. In Sentry, create a project and set its platform to **Node.js** so Convex exception events are processed correctly. Copy the DSN.
2. In the Convex dashboard, open **Settings**, then **Integrations**, click the Sentry card and paste the DSN. You can add tags of your own if you like.
3. Convex tags every event with `func`, `func_type`, `func_runtime`, `request_id`, `server_name` and `environment` (`prod`, `dev` or `preview`). When the function is authenticated, it sets the user to the caller's `tokenIdentifier`. You can't override these tags.
4. Expect exceptions to take a minute or two to show up.

For PostHog Error Tracking, use the PostHog card and your project token. Each `$exception` event carries `convex_function`, `convex_function_type`, `convex_deployment_type` and `convex_request_id`, and stack traces are included. For Datadog Error Tracking, set it up through Datadog's Sentry SDK path, then use the Convex Sentry integration.

## Step 3: Turn function_execution events into alerts

### Axiom monitors in APL

These queries follow the field names Axiom documents for the Convex integration. The first one pages when a function's error rate goes above 5%:

```kusto
['convex']
| where ['data.topic'] == "function_execution"
| where ['convex.deployment_type'] == "prod"
| summarize total_calls = count(), failures = countif(['data.status'] == "failure") by ['data.function.path']
| where total_calls >= 20
| extend error_rate = (failures * 100) / total_calls
| where error_rate > 5
```

The `total_calls` floor stops a function called twice in the window from paging you over one failure. Before you rely on the `deployment_type` filter, check which values your own events carry.

This one finds functions that took longer than 5,000 ms:

```kusto
['convex']
| where ['data.topic'] == "function_execution" and ['data.execution_time_ms'] > 5000
| summarize count() by ['data.function.path']
```

This one ranks functions by how much data they read. It is the query to run once a week, and the one to reach for when database I/O costs climb:

```kusto
['convex']
| where ['data.topic'] == "function_execution"
| summarize avg_docs_read = avg(['data.usage.database_read_documents']), avg_bytes_read = avg(['data.usage.database_io_read_bytes']) by ['data.function.path']
| order by avg_bytes_read desc
```

Axiom's own example uses `database_read_bytes`, but the current Convex event schema calls this field `database_io_read_bytes`, so the query above uses that name.

Axiom monitors can alert on the average execution time too. For example, you can fire when the average across functions goes above 2 seconds in any 5 minute period, and route the alert to email, Slack, PagerDuty or another notifier.

### Datadog and PostHog

In Datadog, write log monitors on the same fields: failed `function_execution` events grouped by function path, and execution time above your threshold. Use the tags you set in `ddtags` to keep production separate from everything else. In PostHog, the full event is the log record body, and the deployment metadata comes as resource attributes, so you filter on those.

### A webhook receiver of your own

A webhook stream fits when you want the events in your own pipeline or warehouse. Every request is signed. Convex computes an HMAC-SHA256 of the body, encodes it as lowercase hex with a `sha256=` prefix, and sends it in the `x-webhook-signature` header. You'll find the secret in the dashboard when you set up the webhook. Verify the signature in constant time, and reject events whose timestamp is too old:

```typescript
export default {
  async fetch(req: Request, env: { WEBHOOK_SECRET: string }) {
    const body = await req.arrayBuffer();
    const signature = req.headers.get("x-webhook-signature");
    if (!signature) return new Response("Unauthorized", { status: 401 });

    const key = await crypto.subtle.importKey(
      "raw",
      new TextEncoder().encode(env.WEBHOOK_SECRET),
      { name: "HMAC", hash: "SHA-256" },
      false,
      ["verify"],
    );
    const valid = await crypto.subtle.verify(
      "HMAC",
      key,
      Uint8Array.fromHex(signature.replace("sha256=", "")),
      body,
    );
    if (!valid) return new Response("Unauthorized", { status: 401 });

    const events = JSON.parse(new TextDecoder().decode(body));
    if (events[0].timestamp < Date.now() - 5 * 60 * 1000) {
      return new Response("Request expired", { status: 403 });
    }

    const failures = events.filter(
      (e: any) => e.topic === "function_execution" && e.status === "failure",
    );
    // forward failures to your alerting tool here
    return new Response("ok");
  },
};
```

## Reading write conflicts, retries and query reruns

Some Convex problems never show up as a failed request. Three fields catch them.

**`occ_info`** appears when a function hits a write conflict with another function. It names the `table_name`, the `document_id`, the `write_source` (the conflicting function) and a `retry_count`. When the same document keeps showing up, many callers are writing to one hot document. A counter that every request increments is the usual culprit. The Health page and `npx convex insights` list the same conflicts.

**`mutation_retry_count`** counts failed attempts before a mutation succeeded. The mutation reports `success`, so an error-rate alert never sees it. Users still wait for every retry. Track the average retry count per mutation next to latency.

**`run_reason`** says why a function ran: `initialSubscription`, `dataChange`, `identityChange`, `webSocket`, `httpApi`, `httpEndpoint`, `cron`, `scheduler`, `action` or `tester`. A query with a high count of `dataChange` reruns is being invalidated by frequent writes to data it reads. Subscription updates count as function calls in billing, so this field explains a function call bill that grows faster than your traffic. `mutation_queue_length` does the same job for mutation backlogs within a single client session.

## Catching functions before they hit Convex limits

Convex enforces limits that turn a slow-growing problem into a sudden failure. The ones worth knowing:

- Actions on the Convex runtime get 64 MiB of RAM, and Node.js actions get 512 MiB. Watch `action_memory_used_mb` for actions creeping toward their runtime's limit.
- Function arguments and return values can each be up to 16 MiB. Node actions accept arguments up to 5MiB. The `function_args_bytes` and `function_returns_bytes` usage fields show how close you are.
- A document can be up to 1 MiB, and an array can hold up to 8192 elements. A list you keep appending to inside one document will hit one of these limits eventually.
- Code size is capped at 32 MiB per deployment.
- The Free plan includes 1,000,000 function calls a month in total. Professional includes 25,000,000 a month, and calls beyond that are billed.

Convex warns you before some of these fail: it adds a `console` event with a `system_code` when a function gets close to a limit. Alert on any `console` event where `system_code` is present. The Health page insights flag functions that read too many bytes or documents in one transaction.

## Following one failed request from the browser to the log line

In production, Convex hides server-side error details from the browser console so it doesn't leak server state. What the client gets is a Request ID. Tie your tools together through it:

1. Log the Request ID in your frontend's error handler, or pass it to your client-side error tracker.
2. For a recent failure, paste the ID into the search box on the Logs page. You'll see every line from that request, including the exception.
3. For an older failure, search your log stream for `function.request_id` (under `data.` in Axiom). In Sentry, search the `request_id` tag.
4. With the failing function identified, read the relevant data using `npx convex run --prod --inline-query` to check whether the documents look the way the code expects.
5. Compare the failure's start time with **Last deployed** on the Health page, and with the deploy history in your CI.

## Where Convex monitoring goes wrong

- **Treating the Logs page as your history.** It keeps a short window, and the request you need will have gone by the next morning. Fix: set up a log stream before you launch.
- **Streaming in the legacy format.** Log streams set up before May 23, 2024 use the legacy event schema, which has different field names. Fix: update the stream to the new format in the dashboard under Settings, Integrations, then update your queries.
- **Letting dev and preview deployments page you.** Exception tags and event metadata cover every deployment type. Fix: filter alerts on `environment` in Sentry and on `convex.deployment_type` in your log tool.
- **Accepting unsigned webhook traffic.** Anyone who learns the URL can post fake events to it. Fix: verify `x-webhook-signature` in constant time and reject stale timestamps.
- **Sentry events arriving garbled.** The Sentry project platform wasn't set to Node.js. Fix: create the project with the Node.js platform.
- **Error-rate alerts that fire on one failure.** A function with two calls and one failure has a 50% error rate. Fix: require a minimum number of calls in the window before the rate counts.
- **Missing slow mutations that succeed.** Write conflicts get retried, and the final `success` hides the cost. Fix: alert on `occ_info` and on the average `mutation_retry_count` per mutation.
- **Writing queries against the wrong field paths.** Axiom nests event data under `data`, while webhook payloads don't. Fix: look at one raw event in your destination before you write a monitor.

## Triaging Convex alerts with Polylane

If you connect a Convex deployment as a cloud account in Polylane, its account page has an **Alerts** tab that triages the account's alerts automatically. Each alert that fires becomes an issue. A confirmed issue gets a fix run, which traces the cause and, when a code change fixes it, opens a pull request in your repository for you to review.

Running on Convex? See [how Polylane monitors Convex in production](https://polylane.com/for/convex/).

## Common questions

**Do I need a paid Convex plan to monitor production?**

No for the basics. The Health page, the Logs page, npx convex logs and npx convex insights don't need a log stream or exception integration. Log streams to Axiom, Datadog, PostHog or a webhook, and exception reporting to Sentry or PostHog, both need a Convex Pro plan.

**How long does the Convex dashboard keep logs?**

The Logs page shows a short history of recent function logs and new ones as they arrive. It isn't a complete historical record, so older requests may be missing. To keep history, set up a log stream, or run npx convex logs --prod | tee ./logs.txt to save a copy locally.

**Which log and error destinations does Convex support?**

Log streams can go to Axiom, Datadog, PostHog, or a webhook at any URL you choose. Exception reporting supports Sentry and PostHog Error Tracking. Datadog Error Tracking works through its Sentry SDK path combined with the Convex Sentry integration.

**How do I find the server error behind a failure a user saw in production?**

Convex hides server error details from the browser in production and gives the client a Request ID. Paste that ID into the search box on the Logs page for recent requests. For older ones, search function.request_id in your log stream or the request_id tag in Sentry.

**What should my first Convex alert be?**

Start with a failure-rate alert per function: count function_execution events with status failure, divide by total calls, group by function path, and fire above 5%. Set a minimum call count so low-traffic functions can't page you over a single error. Add a slow-function alert on execution_time_ms next.

**How do I monitor write conflicts in Convex?**

Failed or retried functions carry an occ_info object with the table name, document ID and the conflicting function. Alert on events where it's present, and track mutation_retry_count on successful mutations. Run npx convex insights --prod --details to see conflicts from the last 72 hours.

**Why does my Convex log stream use different field names from the docs?**

Log streams configured before May 23, 2024 send the legacy event schema. Convex recommends updating to the current format. Once you do, update your queries and monitors to the new field names.

**How quickly do Convex exceptions appear in Sentry or PostHog?**

Convex documents a delay of a minute or two for both Sentry and PostHog. For anything faster, alert from the log stream on failed function_execution events, which carry the error message and stack trace.

## Sources

- [Log Streams, Convex Developer Hub](https://docs.convex.dev/production/integrations/log-streams/)
- [Legacy Event Schema, Convex Developer Hub](https://docs.convex.dev/production/integrations/log-streams/legacy-event-schema)
- [Exception Reporting, Convex Developer Hub](https://docs.convex.dev/production/integrations/exception-reporting)
- [Health, Convex Developer Hub](https://docs.convex.dev/dashboard/deployments/health)
- [Logs, Convex Developer Hub](https://docs.convex.dev/dashboard/deployments/logs)
- [CLI, Convex Developer Hub](https://docs.convex.dev/cli)
- [Limits, Convex Developer Hub](https://docs.convex.dev/production/state/limits#functions)
- [Observing your app in production, Stack by Convex](https://stack.convex.dev/observability-in-production)
- [Connect Axiom with Convex, Axiom Docs](https://axiom.co/docs/apps/convex)
- [Send data from Convex to Axiom, Axiom Docs](https://axiom.co/docs/send-data/convex)
- [Convex + Axiom: Complete observability for reactive backends](https://axiom.co/blog/axiom-convex-integration)
- [Polylane documentation](https://docs.polylane.com/llms-full.txt)

## About the author

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

## More in this series

[Platform playbooks](https://polylane.com/series/platform-playbooks/)

1. [How to Monitor a Django App on Render](https://polylane.com/learn/monitoring-coverage/how-to-monitor-a-django-app-on-render/)
2. [How to Monitor a FastAPI App on Railway: Logs, Traces and Alerts](https://polylane.com/learn/monitoring-coverage/how-to-monitor-a-fastapi-app-on-railway/)
3. [How to Monitor a Supabase App in Production](https://polylane.com/learn/monitoring-coverage/how-to-monitor-a-supabase-app-in-production/)
4. [How to Debug Cloudflare Workers Errors: Logs, Traces and Error Codes](https://polylane.com/learn/troubleshooting/how-to-debug-cloudflare-workers-errors/)
5. [How to debug Vercel function timeouts](https://polylane.com/learn/troubleshooting/how-to-debug-vercel-function-timeouts/)
6. [Vercel 504 Gateway Timeout on Serverless Functions: Causes and Fixes](https://polylane.com/learn/troubleshooting/vercel-504-gateway-timeout-on-serverless-functions/)
7. [Cloudflare Workers error 1101: causes and how to fix it](https://polylane.com/learn/troubleshooting/cloudflare-workers-error-1101/)
8. [Cloudflare Hyperdrive connection errors: causes and fixes](https://polylane.com/learn/troubleshooting/cloudflare-hyperdrive-connection-errors/)
9. [How to Monitor a Convex App in Production](https://polylane.com/learn/monitoring-coverage/how-to-monitor-a-convex-app-in-production/)

Previous: [Cloudflare Hyperdrive connection errors: causes and fixes](https://polylane.com/learn/troubleshooting/cloudflare-hyperdrive-connection-errors/)

## Related

- [How to Use Claude Code to Debug Production Issues](https://polylane.com/learn/ai-in-production/how-to-use-claude-code-to-debug-production-issues/): Debug production issues with Claude Code: get logs and traces into the session, test hypotheses against evidence, stay read-only and ship a verified fix.
- [How to monitor a vibe-coded app in production](https://polylane.com/learn/monitoring-coverage/vibe-coding-production-monitoring/): Monitor a vibe-coded app in production: surface swallowed errors, add a health route, tag deploys, watch cron jobs and alert on what users feel first.

Get started with one command: `curl -fsSL https://polylane.com/setup | bash` installs the CLI, connects your coding agents, and creates the account.
