Get started Dashboard
Monitoring coverage ·
Part 2 of Platform playbooks

How to Monitor a FastAPI App on Railway: Logs, Traces and Alerts

Explore with AI

Give the FastAPI app a /health route that Railway checks on each deploy, write single-line JSON logs to stdout, and turn on tracing so Uvicorn runs under opentelemetry-instrument and requests show up on Railway's Traces page. Then add monitors on memory and CPU in the Observability dashboard (Pro plan), plus a project webhook that posts failed and crashed deploys to Slack or Discord. Railway does not collect request latency or error-rate metrics, and it never runs the healthcheck again after a deploy goes live, so use its HTTP logs or traces for latency and an outside uptime check for availability.

On this page

Monitoring a FastAPI app on Railway comes down to a few pieces: a /health route that Railway checks on every deploy, JSON logs written to stdout, OpenTelemetry traces sent to Railway’s built-in collector, monitors on CPU and memory, and a project webhook that posts failed and crashed deploys to Slack or Discord. Railway graphs CPU, memory, disk and network for every service with no setup. It does not collect request latency or error rates as metrics, and it only calls your health check while a deploy goes live.

You add very little code to the app: a health route, a log formatter and one middleware. Tracing needs a package install and a new start command. The alerts are dashboard settings.

What monitoring means for a FastAPI service on Railway

Here, monitoring means you can answer these questions without opening a shell: is the service up, is it healthy under load, what did a slow or failed request do, and who hears about it when something breaks. Railway answers each one with a different feature, and each feature has its own limits.

  • Platform metrics. Every service has a Metrics tab with CPU, memory, disk usage and network traffic. It keeps up to 30 days of data, and dotted lines mark each deploy, so you can see which commit started a memory climb. With several replicas, the Sum view adds them together and the Replica view shows each one separately.
  • Logs. Railway captures anything your process writes to stdout or stderr, plus build and deploy logs. Railway’s edge also writes an HTTP log for every request, with the path, status code and timings, and you can filter on all of them.
  • Traces. For a traced service, Railway’s edge starts a trace on each request and passes a W3C traceparent header to your app. If the app runs an OpenTelemetry SDK, its spans join the same trace on the Traces page. Tracing is a preview feature.
  • Alerts. Project webhooks fire when a deploy changes state, including Failed and Crashed, and when a monitor alerts. Monitors on the Observability dashboard alert when CPU, RAM, disk or network egress crosses a threshold. Monitors need the Pro plan.

What Railway leaves to you

For a Python API, these gaps matter. Railway does not collect application metrics such as request latency, error rate or business counts. Its docs send you to a third-party tool for those. The healthcheck only runs at the start of a deploy, and Railway never uses it to keep watching the service. There is also no log drain setting, so sending logs somewhere else takes a forwarder or an SDK. The sections below close each gap with what Railway already gives you, then show when an outside tool is worth adding.

Add a health route and bind Uvicorn to Railway’s port

Railway injects a PORT variable and uses it for health checks, so Uvicorn has to listen on 0.0.0.0 and $PORT. Add a cheap route that returns 200 once the app can serve requests:

# main.py
from fastapi import FastAPI

app = FastAPI()

@app.get("/health")
async def health():
    return {"status": "ok"}

Then set the health check path and start command. You can do this in the service’s Settings, or in a railway.toml next to your code:

[deploy]
startCommand = "uvicorn main:app --host 0.0.0.0 --port $PORT --no-access-log"
healthcheckPath = "/health"
healthcheckTimeout = 300
restartPolicyType = "ON_FAILURE"

On each deploy, Railway calls /health until it gets a 2xx response, then moves traffic to the new version. If no 2xx arrives within the timeout, which is 300 seconds by default, Railway marks the deploy as failed. To give the app longer, for example when it loads a large model at startup, set the RAILWAY_HEALTHCHECK_TIMEOUT_SEC variable.

These details trip up FastAPI apps:

  1. The check comes from the hostname healthcheck.railway.app. If you use TrustedHostMiddleware, add that host to the allowed list, or the check fails with a 400:
from fastapi.middleware.trustedhost import TrustedHostMiddleware

app.add_middleware(
    TrustedHostMiddleware,
    allowed_hosts=["api.example.com", "*.up.railway.app", "healthcheck.railway.app"],
)
  1. Keep /health shallow. If the route also queries Postgres, a short database blip turns into a failed deploy. Check dependencies from a separate route and watch that one from outside Railway (see the last setup section).

Railway has deprecated config-as-code files in favour of Infrastructure as Code. Existing railway.toml files keep working for legacy services until 2026-12-01. If you’d prefer not to keep a file, the same settings are in the dashboard.

Write one JSON log line per request to stdout

Railway treats a log line as structured when the whole JSON object sits on one line. It reads message and level, colours each line by level, and lets you filter on any other key as @key:value. Plain text lines are still stored. They get a level from their stream, info for stdout and error for stderr, so @level filters work on them, but they carry no custom attributes to filter on.

Python’s logging.StreamHandler writes to stderr unless you tell it otherwise, and Railway converts every stderr line to level: error. Point the handler at stdout and healthy requests stop showing up in red:

# logging_setup.py
import json
import logging
import sys

class JsonFormatter(logging.Formatter):
    def format(self, record):
        payload = {
            "message": record.getMessage(),
            "level": record.levelname.lower(),
            "logger": record.name,
        }
        payload.update(getattr(record, "fields", {}))
        if record.exc_info:
            payload["exception"] = self.formatException(record.exc_info)
        return json.dumps(payload, default=str)

def setup_logging():
    handler = logging.StreamHandler(sys.stdout)
    handler.setFormatter(JsonFormatter())
    logging.basicConfig(level=logging.INFO, handlers=[handler], force=True)

The stack trace goes into a single exception field, so a crash stays as one entry you can open.

Next, log one event per request with the route, status, duration and trace ID:

# main.py
import logging
import time

from fastapi import FastAPI, Request
from opentelemetry import trace

from logging_setup import setup_logging

setup_logging()
log = logging.getLogger("app")
app = FastAPI()

@app.middleware("http")
async def log_requests(request: Request, call_next):
    start = time.perf_counter()
    fields = {"method": request.method, "path": request.url.path}
    try:
        response = await call_next(request)
    except Exception:
        fields["duration_ms"] = round((time.perf_counter() - start) * 1000, 1)
        log.exception("unhandled error", extra={"fields": fields})
        raise
    ctx = trace.get_current_span().get_span_context()
    fields.update(
        status_code=response.status_code,
        duration_ms=round((time.perf_counter() - start) * 1000, 1),
        trace_id=format(ctx.trace_id, "032x") if ctx.is_valid else None,
    )
    log.info("request", extra={"fields": fields})
    return response

trace_id stays empty until you turn on tracing in the next sections. After that, every log line links to its trace.

I built observability tools at Baselime and then led the Workers observability team at Cloudflare, and this is the habit I push hardest: write one wide event per request, with enough fields that you never need a second query to explain it.

The --no-access-log flag in the start command matters. Railway’s HTTP logs already record every request at the edge, and Uvicorn’s plain-text access lines only eat into the logging limit of 500 lines per second per replica. Above that limit, Railway drops lines and prints a warning with the number dropped.

Once the app is deployed, these queries work in the Log Explorer (the Observability tab) and in each deployment’s log panel:

@level:error
@path:/orders AND @status_code:>=500
@duration_ms:>500
@level:error AND "unhandled error"

In a terminal, railway logs shows the latest deployment’s logs.

Your plan sets how long Railway keeps logs, and traces are kept for the same period. If you upgrade, logs that had aged out come back straight away.

Railway log and trace retention by plan
PlanRetention
Free 3 days
Trial 7 days
Hobby 7 days
Pro 30 days
Enterprise Up to 90 days

Read latency and error rates from Railway’s HTTP logs

Before you write any metrics code, check the edge’s HTTP logs. They answer most latency and error questions. Each request has @httpStatus, @responseTime (time to first byte), @totalDuration, @path, @method and @srcIp, and the number fields accept comparisons and ranges:

@httpStatus:500..599
@path:/api/v1/orders AND @httpStatus:>=400
@responseTime:>1000
@totalDuration:>5000 @httpStatus:>=500
-@httpStatus:200

Railway only fills @responseDetails when the application fails to respond, so it explains requests that never reached a FastAPI handler. These filters give you error counts and slow requests per route, which covers most of what a small team asks. They can’t show where the time went inside the handler. Traces do that.

Turn on tracing and let OpenTelemetry instrument FastAPI

You turn tracing on per service and per environment. Once it’s on, the edge traces every request to the service, and Railway adds the OTEL_* variables on the next deploy. There is no sample rate: Railway traces every request.

  1. Enable tracing for the service in the Tracing setup panel on the Traces tab, or from the CLI:
railway trace enable --service api
railway trace status --all
  1. Install the OpenTelemetry distribution and let it detect FastAPI and your other libraries:
pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install

opentelemetry-bootstrap installs the matching packages, such as opentelemetry-instrumentation-fastapi. Run it locally, then add what it installed to requirements.txt or pyproject.toml so the Railway build installs them too.

  1. Put opentelemetry-instrument in front of the start command:
opentelemetry-instrument uvicorn main:app --host 0.0.0.0 --port $PORT --no-access-log
  1. Railway’s receiver only accepts traces, but the Python distribution exports metrics by default. Add these variables to the service:
OTEL_METRICS_EXPORTER=none
OTEL_LOGS_EXPORTER=none
  1. Redeploy, send a few requests, and wait for the App indicator in the Tracing setup panel to turn green.

Each request now produces a server span named after its method and route, with http.route and http.response.status_code. Calls made with httpx, requests or aiohttp become client spans. asyncpg, psycopg2, SQLAlchemy and redis produce spans that include the statement or command. To see your own slow work in the waterfall, wrap it in a span:

from opentelemetry import trace

tracer = trace.get_tracer("checkout")

def calculate_total(cart):
    with tracer.start_as_current_span("calculate-total") as span:
        span.set_attribute("cart.items", len(cart.items))
        return price_items(cart)

The Traces page and the CLI use the same filter syntax, so you can find slow routes and errors from either one:

railway trace list --since 30m --errors
railway trace list --all --filter '@http.route:/checkout @duration:>500'

When a user reports a failed request, ask them for the x-railway-trace-id response header, or include it in your error responses. Paste it into the Trace ID field to open that exact request.

To skip the SDK, use the Automatic instrumentation switch. It traces Python processes with no code changes, on a best-effort basis. The SDK gives you custom spans and full control.

Tracing has limits. Each replica can export 1,000 spans per 10 seconds, and a search returns up to 500 traces. On a busy service, set OTEL_TRACES_SAMPLER=traceidratio and put a fraction in OTEL_TRACES_SAMPLER_ARG. The default parent-based sampler follows the edge, which traces everything.

Put the numbers on one dashboard and add monitors

  1. Open the Observability tab and check that the environment selector shows production. The dashboard only covers one environment.
  2. Click Start with a simple dashboard. Railway creates widgets for spend, service metrics and logs.
  3. Add a logs widget for the API service filtered to @level:error, so errors sit next to the memory graph.
  4. On the memory widget, open the three-dot menu, choose Add monitor, and set a threshold above your normal peak. Do the same for CPU.
  5. If the service has a volume, add a disk usage monitor. Add a network egress monitor if you watch bandwidth costs.

Monitors send email and in-app notifications. They also fire your project webhook, so they arrive in the same channel as failed deploys. For a Python service, the memory monitor matters most: a slow leak shows up on the graph well before the process crashes, and the alert gives you time to act.

Send failed and crashed deploys to Slack or Discord

You set webhooks per project, and they fire for every environment in that project.

  1. Create an incoming webhook in Slack (a hooks.slack.com URL), or in a Discord channel under Integrations, then Webhooks.
  2. In the Railway project, open Settings, then the Webhooks tab. Paste the URL, choose the events and click Save Webhook.
  3. Trigger a real deploy to check it works. Test Webhook sends from your browser, so CORS on the receiving end can make a working endpoint look broken.

Railway recognises Slack and Discord URLs and reformats the payload for them. Any other URL gets JSON with a type such as Deployment.failed and a resource block naming the project, environment and service. That’s enough for a small receiver to keep only production events. Host the receiver in its own project, so an outage in the app can’t take its alerting down with it.

By default, Railway restarts a crashed service on failure, up to 10 times, so a one-off crash often fixes itself before anyone looks. A failed deploy never fixes itself. Treat Deployment.failed as urgent, and page someone when the same service crashes repeatedly.

Cover the checks Railway stops running after deploy

The healthcheck stops once a deploy is live, so nothing on Railway tells you when the API stops answering overnight. For continuous checks, Railway’s docs point to the Uptime Kuma template in its marketplace. Run it in a separate project and point it at a route that checks your real dependencies.

For longer retention or application metrics, send telemetry to a third-party tool. Either path works:

  • Vendor SDK. For Python, Datadog wraps the start command with ddtrace-run, New Relic uses newrelic-admin run-program, and you initialise Sentry in code. Map Railway’s variables to the vendor’s names, for example DD_SERVICE=${{RAILWAY_SERVICE_NAME}} and DD_VERSION=${{RAILWAY_DEPLOYMENT_ID}}.
  • OpenTelemetry to your own backend. Set OTEL_EXPORTER_OTLP_ENDPOINT to the backend, and prefer http/protobuf on port 4318. If you set this variable yourself, Railway adds none of its tracing variables, so your app’s spans no longer reach the Traces page. The edge keeps tracing requests.

To forward raw stdout, run Vector or Fluent Bit as its own Railway service.

Where FastAPI monitoring on Railway goes wrong

  • Uvicorn is not listening on the PORT variable. The health check returns service unavailable and the deploy fails. Start Uvicorn with --host 0.0.0.0 --port $PORT.
  • Every log line is red. The handler writes to stderr. Pass sys.stdout to StreamHandler.
  • JSON logs show up as plain text. Pretty-printed or multi-line JSON breaks parsing. Write one object per line with json.dumps and no indent.
  • The health check fails with status 400. TrustedHostMiddleware is rejecting healthcheck.railway.app. Add it to allowed_hosts.
  • Edge spans appear but the app’s spans don’t. The OTEL_* variables arrive on the first deploy after you enable tracing, so redeploy. Also check that the service doesn’t set its own OTEL_EXPORTER_OTLP_ENDPOINT.
  • The SDK logs export errors. It’s trying to send metrics or logs. Set OTEL_METRICS_EXPORTER=none and OTEL_LOGS_EXPORTER=none.
  • Staging shows nothing. Tracing, the dashboard and monitors are all per environment. Turn them on in each environment.
  • Logs go missing under load. You’ve hit 500 lines per second per replica. Turn off Uvicorn access logs, raise the log level in production, and sample high-frequency events.

Go-live checklist for FastAPI on Railway

  • /health returns 200 and is set as the healthcheck path
  • Uvicorn listens on 0.0.0.0 and $PORT
  • Logs are single-line JSON on stdout with level, path, status_code, duration_ms and trace_id
  • Tracing is on in production and the App indicator is green
  • OTEL_METRICS_EXPORTER and OTEL_LOGS_EXPORTER are both set to none
  • The production environment has a dashboard with an error log widget
  • Memory and CPU monitors are set up (Pro plan)
  • A project webhook posts to your alerts channel and has fired on a real deploy
  • An outside uptime check hits a route that checks your dependencies

Let Polylane watch the Railway project for you

Polylane connects to Railway with OAuth or a workspace token. It syncs your projects, services and deployments, watches deploys, reads logs and checks service metrics for issues, with no alert rules to write or thresholds to tune. Railway’s logs and metrics are enough on their own as a telemetry source, and each confirmed issue goes to one agent that traces the cause and opens the fix as a pull request for you to review.

Running on Railway? See how Polylane monitors Railway in production.

Common questions.

Does Railway track request latency and error rates for a FastAPI app?

Not as metrics. The Metrics tab only covers CPU, memory, disk and network. Railway's HTTP logs do record `@httpStatus`, `@responseTime` and `@totalDuration` for every request, so filters like `@httpStatus:500..599` or `@responseTime:>1000` find errors and slow requests. For timing inside each route, enable tracing and run the app under `opentelemetry-instrument`.

Is Railway's healthcheck enough to know my API is up?

No. Railway only calls the healthcheck path at the start of a deploy, until it gets a 2xx, and never after the deploy is live. For continuous checks, Railway's docs suggest the Uptime Kuma template, or you can point any outside uptime service at your public domain.

How long does Railway keep logs, traces and metrics?

Railway keeps logs for 3 days on Free, 7 days on Trial and Hobby, 30 days on Pro and up to 90 days on Enterprise. Traces are kept for the same period. Metrics graphs hold up to 30 days of data. For longer history, send telemetry to a third-party tool with a vendor SDK or OpenTelemetry.

Why are all my FastAPI logs marked as errors on Railway?

By default, Python's logging StreamHandler writes to stderr, and Railway converts every stderr line to level error. Create the handler with `logging.StreamHandler(sys.stdout)` and include a `level` field in single-line JSON. Railway then colours each line by its real level.

Do I need the Pro plan to get alerts?

Monitors need the Pro plan. They are the threshold alerts on CPU, RAM, disk usage and network egress. Project webhooks cover failed builds, failed deploys and crashed deployments, and Railway formats them automatically for Slack and Discord URLs.

Can I send FastAPI traces to Datadog or Grafana and keep Railway's Traces page?

Not from the same exporter. If you set `OTEL_EXPORTER_OTLP_ENDPOINT` yourself, Railway adds none of its tracing variables, and your app's spans stop reaching the Traces page. The edge still traces requests. Choose one destination per service: Railway's built-in tracing, or a vendor setup such as `ddtrace-run` or an OTLP backend.

What does the Railway rate limit of 500 logs/sec warning mean?

Railway allows 500 log lines per second per replica on every plan and drops anything above that. For FastAPI, start Uvicorn with `--no-access-log`, because the edge HTTP logs already record every request. Then raise the production log level and sample high-frequency events.

How do I find the trace for one user's failed request?

Railway's edge adds an `x-railway-trace-id` header to every traced response. Have the client or your error page show it, then paste it into the Trace ID field on the Traces page, or run `railway trace get` with the ID.

Sources

  1. Railway Observability Dashboard
  2. Railway Logs
  3. Railway Metrics
  4. Railway Healthchecks
  5. Railway Tracing
  6. Railway Tracing for Python
  7. Set Up Alerts for Crashes, Restarts, and Failed Deploys
  8. Connect a Third-Party Observability Tool
  9. Railway Config as Code reference
  10. Product docs: Quickstart, Clouds and Detect
  11. Product overview and build log

About the author

Boris Tane

Founder of Polylane

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

More in this series

Platform playbooks

  1. 1 How to Monitor a Django App on Render
  2. 2 How to Monitor a FastAPI App on Railway: Logs, Traces and Alerts
  3. 3 How to Monitor a Supabase App in Production
  4. 4 How to Debug Cloudflare Workers Errors: Logs, Traces and Error Codes
  5. 5 How to debug Vercel function timeouts
  6. 6 Vercel 504 Gateway Timeout on Serverless Functions: Causes and Fixes
  7. 7 Cloudflare Workers error 1101: causes and how to fix it
  8. 8 Cloudflare Hyperdrive connection errors: causes and fixes
  9. 9 How to Monitor a Convex App in Production

Related

Nobody should be on-call. Polylane watches your infra, finds what broke, and writes the fix.

Get started for free