Get started Dashboard
Monitoring coverage ·
Part 1 of Platform playbooks

How to Monitor a Django App on Render

Explore with AI

Start with what Render already collects: logs from stdout and stderr on every service, CPU, memory and HTTP metrics, and health checks on the web service. Then log JSON with Render's Rndr-Id request ID, point an HTTP health check at a Django view that queries the database, turn on failure notifications, and stream logs and metrics out before retention ends. Add OpenTelemetry tracing set up in a Gunicorn post_fork hook, and heartbeats for Celery tasks and cron jobs, because health checks never reach them.

On this page

To monitor a Django app on Render, start with what the platform already collects, then close the gaps it leaves. Render captures everything your web service, Celery worker and cron jobs write to stdout and stderr. It charts CPU, memory and HTTP traffic, and it probes web services with health checks. Your job is to make those logs searchable by request, give the health check a path that proves Django can reach its database, turn on failure notifications, and stream logs and metrics to a tool that keeps them longer than Render does.

Two things stay invisible until you add them: what happens inside a slow request, and whether a background job finished its work. For the first, add OpenTelemetry tracing to Django, set up so it works with the Gunicorn --preload flag Render sets by default. For the second, give each job that matters a heartbeat. The steps below take a typical project (a Gunicorn web service, a Celery worker, a cron job and Render Postgres) from default settings to alerts that reach you.

What Render already watches on each Django service

Each part of a Django deployment runs as its own Render service, and each one gets different signals:

  • Web service running Gunicorn and Django. Runtime logs, CPU and memory, HTTP request volume from the public internet, and a health check every few seconds. A Pro workspace or higher adds a log line for every public HTTP request and a Response Times graph at p50, p75, p90 and p99.
  • Background worker running Celery. Runtime logs, CPU and memory. Health checks only cover web services and private services, so nothing probes a worker.
  • Cron job. Logs and run history on its Runs page, plus a notification when a run fails. Render doesn’t emit streamed metrics for cron jobs.
  • Render Postgres. Active connections, transaction volume, replication lag once you add a read replica, lock-delayed queries, running processes and top queries.

None of this tells you which view is slow and why, or whether the task that sends receipts actually ran. The steps below cover both.

Step 1: Log JSON to stdout with Render’s request ID on every line

Render needs no SDK for logs, because it reads stdout and stderr. What it can’t do is link your log lines to the request that produced them. For every public request, Render writes a requestID into its HTTP request log. It sends the same value to your service in the Rndr-Id header and returns it to the client in the response. Cloudflare also sets a CF-Ray header on every inbound request, and Render’s uptime guide recommends logging it. I led the Workers observability team at Cloudflare, and on any platform I’d make this the first change: put the platform’s request ID into your own log lines.

Store both IDs in context variables, so every log line written during the request picks them up:

# myproject/log_context.py
import contextvars
import json
import logging
import os

request_id = contextvars.ContextVar("request_id", default="-")
cf_ray = contextvars.ContextVar("cf_ray", default="-")

COMMIT = os.environ.get("RENDER_GIT_COMMIT", "")[:7]
INSTANCE = os.environ.get("RENDER_INSTANCE_ID", "")


class JsonFormatter(logging.Formatter):
    def format(self, record):
        payload = {
            "level": record.levelname.lower(),
            "logger": record.name,
            "message": record.getMessage(),
            "requestId": request_id.get(),
            "cfRay": cf_ray.get(),
            "commit": COMMIT,
            "instance": INSTANCE,
        }
        if record.exc_info:
            payload["exception"] = self.formatException(record.exc_info)
        return json.dumps(payload)


class RequestIdMiddleware:
    def __init__(self, get_response):
        self.get_response = get_response

    def __call__(self, request):
        rid = request_id.set(request.headers.get("Rndr-Id", "-"))
        ray = cf_ray.set(request.headers.get("CF-Ray", "-"))
        try:
            return self.get_response(request)
        finally:
            request_id.reset(rid)
            cf_ray.reset(ray)

Put the middleware first in MIDDLEWARE and send Django’s logging to stdout:

# settings.py
MIDDLEWARE = [
    "myproject.log_context.RequestIdMiddleware",
    "django.middleware.security.SecurityMiddleware",
    # ...
]

LOGGING = {
    "version": 1,
    "disable_existing_loggers": False,
    "formatters": {"json": {"()": "myproject.log_context.JsonFormatter"}},
    "handlers": {
        "stdout": {
            "class": "logging.StreamHandler",
            "stream": "ext://sys.stdout",
            "formatter": "json",
        },
    },
    "root": {"handlers": ["stdout"], "level": os.environ.get("LOG_LEVEL", "INFO")},
}

Render sets RENDER_GIT_COMMIT and RENDER_INSTANCE_ID on every service by default, so each line now records which deploy and which instance wrote it. Gunicorn’s access log needs no setup. Render’s Python runtime sets GUNICORN_CMD_ARGS to --preload --access-logfile - --bind=0.0.0.0:10000, which already sends access lines to stdout.

The log explorer can now answer questions like these. Its search box accepts wildcards, and RE2 regular expressions between slashes:

8ebfa3c3-8929-4885            # every line for one request
status_code:/5../             # request logs with a 5xx status
path:api/orders/*             # requests to one endpoint
/responseTimeMS=\d{3}\d+/     # requests slower than one second

To watch lines as they arrive during a deploy, use Live tail in the dashboard or run render logs from the Render CLI.

Step 2: Point the health check at a view that queries the database

By default, Render’s health check is a TCP probe to your service’s port. It passes once Gunicorn accepts a connection, even if every view raises an error. An HTTP health check makes Django prove it can do real work, and Render suggests a simple database query for it.

# myproject/health.py
from django.db import connection
from django.http import JsonResponse


def healthz(request):
    try:
        with connection.cursor() as cursor:
            cursor.execute("SELECT 1")
    except Exception:
        return JsonResponse({"status": "unhealthy"}, status=503)
    return JsonResponse({"status": "ok"})
# urls.py
from myproject.health import healthz

urlpatterns = [
    path("healthz/", healthz),
    # ...
]

Then set the path in either of two places:

  1. In the Render Dashboard, open the web service’s Settings page, scroll to Health Checks, click Edit, enter /healthz/ and click Save Changes.
  2. In your Blueprint, add healthCheckPath to the web service:
services:
  - type: web
    name: django-web
    startCommand: gunicorn myproject.wsgi:application -c gunicorn.conf.py
    healthCheckPath: /healthz/

Check it from your machine. It should print 200:

curl -s -o /dev/null -w "%{http_code}\n" https://your-service.onrender.com/healthz/

What Render does when the check fails

A check passes when the instance returns any 2xx or 3xx status within five seconds. Anything else counts as a failure. If a running instance fails consecutive checks for 15 seconds, Render stops sending it traffic, and any healthy instances keep serving. If the failures last 60 seconds, Render restarts the instance and notifies you. A new deploy only receives traffic once all its new instances pass at the same time. If that hasn’t happened within 15 minutes, Render cancels the deploy and the old version keeps serving.

0 s 15 s 30 s 45 s 60 s Checks failing 15 s Out of routing 45 s
Checks failing: still receives traffic
Out of routing: restarted at 60 s
Figure 1
What happens to a Render instance that keeps failing its health check
From Render's health check docs: routing stops after 15 seconds of failed checks, and the instance restarts at 60 seconds.

Make sure Django accepts the health check’s Host header

Render sets the health check’s Host header to one of your verified custom domains, or to the service’s onrender.com subdomain if you have none. If that host is missing from ALLOWED_HOSTS, Django rejects the request and every deploy fails its checks. Add Render’s hostname from the environment:

ALLOWED_HOSTS = ["www.example.com"]
if os.environ.get("RENDER_EXTERNAL_HOSTNAME"):
    ALLOWED_HOSTS.append(os.environ["RENDER_EXTERNAL_HOSTNAME"])

Keep the view cheap, and check only what this instance needs to serve traffic. A stopped Celery worker or a slow payment API is real trouble, but restarting web instances won’t fix either one, so they belong in Step 7.

Step 3: Turn on failure notifications for every service

From your workspace home, click Integrations > Notifications and set the default to Only failure notifications, by email, Slack or both. At that level, Render tells you when:

  • a build or deploy fails
  • a cron job run fails
  • a running service becomes unhealthy
  • a persistent disk goes past 80% of its allocated space
  • Render suspends a service that keeps failing to start and become healthy over a 24-hour period

All notifications adds Slack messages for deploys that go live and services that recover, which suits a deploy channel. You can override the level for each service, and Render webhooks can send events anywhere else.

Step 4: Read metrics, logs and deploys over the same time range

When a graph on the Metrics page moves, you find the cause fastest by lining up three views over the same window:

  1. Metrics. Check CPU, memory, request volume grouped by status code, and response times. Memory that climbs steadily and ends in restarts suggests a leak. A jump in 5xx responses points at code or a dependency.
  2. Logs. Open the log explorer for that window and filter with status_code:/5../ or the view’s path. Click an instance ID to see only that instance’s lines.
  3. Deploys. Open the service’s Events page, find the deploy just before the change, and read that deploy’s logs. Health check failures during a deploy explain why a new version never went live.

On the database, watch Active Connections and Lock-Delayed Queries. The second counts completed queries that waited one second or longer on a lock. Django keeps a separate database connection for each worker thread, so your database must allow at least as many connections as you have worker threads. Render sets WEB_CONCURRENCY from the plan’s CPU count (1 on 0.5c-512mb, 2 on 2c-4g), so more instances or more workers mean more connections. The Queries tab lists running processes and the most frequent queries.

Watch for two blind spots. HTTP metrics and request logs only include requests from the public internet, so calls between your services over the private network don’t appear. And the outbound bandwidth graph only updates once an hour.

Step 5: Stream logs and metrics out before Render’s retention ends

Render keeps logs and metrics for a period set by your workspace plan. Older logs are deleted, and upgrading later doesn’t bring them back.

Render log and metrics retention by workspace plan
Workspace planLog retentionMetrics retention
Hobby 7 days 7 days
Pro 14 days 14 days
Scale / Enterprise 30 days 30 days

Render also processes at most 6,000 application log lines per minute for each running instance. It drops anything above that from both the explorer and log streams. Keep SQL logging and DEBUG-level logging off in production.

To keep logs longer and alert on them, add a log stream:

  1. From your workspace home, click Integrations > Observability and scroll to Log Streams.
  2. Under Default destination, click + Set default.
  3. Enter your provider’s endpoint: HOST:PORT for syslog, or the full URL for HTTPS.
  4. Add a token if your provider needs one, click Save Changes, and choose whether to include preview instances.

Syslog streams use TLS over TCP in RFC5424 format. On Pro you can leave individual services out of the stream, and on Scale you can send a service to its own destination. A log stream only carries lines. To join them up with other data, use the IDs you logged in Step 1, and have your provider parse requestId and commit as fields.

For metrics, a workspace admin on Pro or higher can open Observability, go to Metrics Stream and click + Add destination. Render lists Better Stack, Datadog, Grafana, Groundcover, Honeycomb, New Relic, Pydantic Logfire and SigNoz, plus a custom OpenTelemetry endpoint. Datadog receives metrics in its own format and needs an organisation-level API key. Neither stream counts against your outbound bandwidth.

Step 6: Trace Django with OpenTelemetry, set up for Gunicorn preload

A trace shows where a slow request spent its time, broken down by view and by query. The opentelemetry-instrumentation-django package instruments Django with a single call to DjangoInstrumentor().instrument().

On Render, the catch is --preload. With preload, Gunicorn loads your app in the master process and then forks the workers. OpenTelemetry’s BatchSpanProcessor exports spans from a background thread, and it is not fork-safe: the child process inherits a lock the parent holds, and it deadlocks. The fix is to instrument Django at import time, then create the tracer provider and span processor in a Gunicorn post_fork hook, which runs in each worker after the fork.

pip install opentelemetry-sdk opentelemetry-instrumentation-django opentelemetry-exporter-otlp
# myproject/wsgi.py
import os

from django.core.wsgi import get_wsgi_application
from opentelemetry.instrumentation.django import DjangoInstrumentor

os.environ.setdefault("DJANGO_SETTINGS_MODULE", "myproject.settings")
DjangoInstrumentor().instrument()
application = get_wsgi_application()
# gunicorn.conf.py
import os

from opentelemetry import trace
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor


def post_fork(server, worker):
    resource = Resource.create(attributes={
        "service.name": os.environ.get("RENDER_SERVICE_NAME", "django-web"),
        "service.version": os.environ.get("RENDER_GIT_COMMIT", "unknown"),
    })
    provider = TracerProvider(resource=resource)
    provider.add_span_processor(
        BatchSpanProcessor(OTLPSpanExporter(endpoint=os.environ["OTEL_EXPORTER_OTLP_ENDPOINT"]))
    )
    trace.set_tracer_provider(provider)

Set OTEL_EXPORTER_OTLP_ENDPOINT, plus any credentials your tracing backend needs, as environment variables on the service. Start Gunicorn with -c gunicorn.conf.py, as in the Blueprint above. Because every span carries the commit SHA, you can compare latency before and after each deploy. To turn off the Django instrumentation without a code change, set OTEL_PYTHON_DJANGO_INSTRUMENT=False.

Step 7: Give Celery tasks and cron jobs a heartbeat

Health checks never reach a background worker, and a cron job that hangs never fails. Render guarantees at most one active run per cron job. A hung run therefore delays every scheduled run behind it until Render stops it after 12 hours. Worker logs and CPU graphs tell you a process is alive. They can’t tell you the job did its work.

Save a timestamp in a shared cache, such as Render Key Value, each time a job succeeds:

# billing/tasks.py
from celery import shared_task
from django.core.cache import cache
from django.utils import timezone


@shared_task
def sync_payments():
    run_payment_sync()  # the real work
    cache.set("heartbeat:sync_payments", timezone.now().timestamp(), timeout=None)

Then expose how old each heartbeat is, and return 503 when one is stale:

# myproject/health.py (continued)
import time

from django.core.cache import cache

MAX_AGE_SECONDS = {"sync_payments": 15 * 60, "nightly_report": 26 * 60 * 60}


def jobs_health(request):
    now = time.time()
    stale = {}
    for job, max_age in MAX_AGE_SECONDS.items():
        last = cache.get(f"heartbeat:{job}")
        if last is None or now - last > max_age:
            stale[job] = None if last is None else int(now - last)
    return JsonResponse({"stale": stale}, status=503 if stale else 200)

Write the heartbeat after the work succeeds. It then proves the whole path from scheduler to finished task. Keep /healthz/jobs/ out of Render’s health check path: a stale heartbeat should page a person, and restarting web instances won’t restart the worker. For cron jobs, make sure the command exits when it finishes. The Runs page shows the logs for each run, and Trigger Run starts one on demand, cancelling any run already in progress.

Step 8: Probe the app from outside Render

Render’s health checks run inside the platform. An external probe sends requests from outside, which is closer to what your users experience, and Render’s uptime guide recommends one. Point it at a real page that reads from the database, at /healthz/jobs/, and at your custom domain, so a broken DNS record or certificate also trips it. When the probe fails, its notification includes the CF-Ray ID, which you can search for in the logs from Step 1. The same guide recommends running more than one instance, so the site stays up while Render moves a service off a failed node.

The alerts worth paging on for Django on Render

  • Deploy failed or service unhealthy, from Render notifications.
  • External probe failing on the main page or the jobs endpoint.
  • 5xx rate rising, built in your log tool from the status codes in request logs.
  • p99 latency above its usual level, from streamed metrics or traces.
  • Memory climbing steadily on web or worker instances.
  • Job heartbeat stale for any task users depend on.
  • Disk past 80%, from Render notifications.
  • Postgres connections near the database’s limit, or a jump in lock-delayed queries.

Page on failed deploys, the external probe, the 5xx rate and stale heartbeats. Send the rest to a channel someone reads every day.

Where Django monitoring on Render goes wrong

  • The health path redirects. Any 3xx counts as healthy, so a path that redirects to HTTPS or to a login page passes even when Django can’t reach the database. Fix: have /healthz/ return 200 or 503 itself, with no auth and no redirect.
  • The custom domain is missing from ALLOWED_HOSTS. Once you verify a domain, Render uses it as the health check’s Host. Fix: list both the custom domain and RENDER_EXTERNAL_HOSTNAME.
  • Tracing starts before the fork. Under --preload, spans stop arriving or workers hang. Fix: create the tracer provider in post_fork.
  • Log lines vanish during an incident. A burst above 6,000 lines per minute per instance gets dropped. Fix: log one structured line per event at INFO, and keep query logging off.
  • Logs expire before anyone reads them. Fix: set up a log stream on day one.
  • The health check calls a third-party API. When that API takes more than five seconds, Render pulls healthy instances out of routing. Fix: check only what the instance needs to serve traffic.
  • Worker failures page nobody. Fix: add heartbeats and an external check on the jobs endpoint.
  • The web service runs one instance. Fix: scale to more than one, so a health check restart doesn’t take the site offline.

Watching Render services and Celery workers with Polylane

Once you connect a Render account, Polylane queries its logs and metrics directly, turns Render alerts into issues, and investigates each one, citing the query or log line behind every claim. Background workers and cron jobs sit in the same graph as your web service and the repository that deploys them, so a regression arrives with the change that landed just before it.

Running on Render? See how Polylane monitors Render in production.

Common questions.

Does Render monitor a Django app without any setup?

Partly. Render captures stdout and stderr from web services, background workers and cron jobs, charts CPU and memory, and probes web services with a TCP health check by default. Its log streams don't create full tracing or APM correlation on their own, and it doesn't check that Celery tasks finished, so add OpenTelemetry and job heartbeats for those.

How long does Render keep logs and metrics?

Both are kept for 7 days on Hobby, 14 days on Pro and 30 days on Scale or Enterprise. Logs older than that are gone even if you upgrade later, so set up a log stream to a provider that keeps them longer.

Why does my Django health check fail on Render?

Most often, Django rejects the Host header because it isn't in ALLOWED_HOSTS. Render sets that header to your verified custom domain, or to the service's onrender.com subdomain if you have none. Other causes are a 4xx or 5xx response from the view, or a response slower than five seconds. On a new deploy, if the new instances don't all pass together within 15 minutes, Render cancels the deploy and keeps the old version running.

How do I monitor Celery workers on Render?

Run Celery as a background worker. That gives you logs and CPU and memory graphs, but no health checks. Have each important task write a timestamp to a shared cache when it succeeds, expose the ages on a web endpoint that returns 503 when one is stale, and point an external monitor at that endpoint.

Which Render monitoring features need a Pro workspace?

HTTP request logs in the log explorer, the Response Times latency graph, metrics streaming to an OpenTelemetry provider, and leaving individual services out of a log stream all need Pro or higher. Sending one service's logs to its own destination needs Scale or higher.

Can I send Render metrics to Datadog or Grafana?

Yes, on Pro or higher. A workspace admin adds a Metrics Stream under Observability. Render pushes metrics to Datadog in Datadog's own format using an organisation-level API key, and to Grafana over OTLP using your stack's endpoint and a token that starts with glc_. Logs go through a separate log stream.

Why do my OpenTelemetry traces stop when Django runs under Gunicorn on Render?

Render's Python runtime starts Gunicorn with --preload, so your app loads before the workers fork. OpenTelemetry's BatchSpanProcessor isn't fork-safe and can deadlock in the child process. Create the tracer provider and span processor in a Gunicorn post_fork hook, and keep DjangoInstrumentor().instrument() in wsgi.py.

Why are some of my log lines missing on Render?

Render processes up to 6,000 application log lines per minute for each instance and drops the excess from both the log explorer and log streams. Turn off DEBUG and SQL query logging in production, and log one structured line per event.

Sources

  1. Logs in the Render Dashboard
  2. How Render handles logging and observability
  3. Service Metrics (Render Docs)
  4. Health Checks (Render Docs)
  5. Streaming Render Service Logs
  6. Email and Slack Notifications (Render Docs)
  7. Streaming Render Service Metrics
  8. Default Environment Variables (Render Docs)
  9. Best Practices for Maximizing Uptime (Render Docs)
  10. Cron Jobs (Render Docs)
  11. Django Instrumentation (OpenTelemetry Python)
  12. Working With Fork Process Models (OpenTelemetry Python)
  13. Databases: persistent connections (Django documentation)
  14. Polylane documentation
  15. Polylane full content

About the author

Boris Tane

Founder of Polylane

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

More in this series

Platform playbooks

  1. 1 How to Monitor a Django App on Render
  2. 2 How to Monitor a FastAPI App on Railway: Logs, Traces and Alerts
  3. 3 How to Monitor a Supabase App in Production
  4. 4 How to Debug Cloudflare Workers Errors: Logs, Traces and Error Codes
  5. 5 How to debug Vercel function timeouts
  6. 6 Vercel 504 Gateway Timeout on Serverless Functions: Causes and Fixes
  7. 7 Cloudflare Workers error 1101: causes and how to fix it
  8. 8 Cloudflare Hyperdrive connection errors: causes and fixes
  9. 9 How to Monitor a Convex App in Production

Related

Nobody should be on-call. Polylane watches your infra, finds what broke, and writes the fix.

Get started for free