How to Monitor a FastAPI App on Railway: Logs, Traces and Alerts
Explore with AI
Give the FastAPI app a /health route that Railway checks on each deploy, write single-line JSON logs to stdout, and turn on tracing so Uvicorn runs under opentelemetry-instrument and requests show up on Railway's Traces page. Then add monitors on memory and CPU in the Observability dashboard (Pro plan), plus a project webhook that posts failed and crashed deploys to Slack or Discord. Railway does not collect request latency or error-rate metrics, and it never runs the healthcheck again after a deploy goes live, so use its HTTP logs or traces for latency and an outside uptime check for availability.
On this page
- What monitoring means for a FastAPI service on Railway
- Add a health route and bind Uvicorn to Railway’s port
- Write one JSON log line per request to stdout
- Read latency and error rates from Railway’s HTTP logs
- Turn on tracing and let OpenTelemetry instrument FastAPI
- Put the numbers on one dashboard and add monitors
- Send failed and crashed deploys to Slack or Discord
- Cover the checks Railway stops running after deploy
- Where FastAPI monitoring on Railway goes wrong
- Go-live checklist for FastAPI on Railway
- Let Polylane watch the Railway project for you
- Common questions
Monitoring a FastAPI app on Railway comes down to a few pieces: a /health route that Railway checks on every deploy, JSON logs written to stdout, OpenTelemetry traces sent to Railway’s built-in collector, monitors on CPU and memory, and a project webhook that posts failed and crashed deploys to Slack or Discord. Railway graphs CPU, memory, disk and network for every service with no setup. It does not collect request latency or error rates as metrics, and it only calls your health check while a deploy goes live.
You add very little code to the app: a health route, a log formatter and one middleware. Tracing needs a package install and a new start command. The alerts are dashboard settings.
What monitoring means for a FastAPI service on Railway
Here, monitoring means you can answer these questions without opening a shell: is the service up, is it healthy under load, what did a slow or failed request do, and who hears about it when something breaks. Railway answers each one with a different feature, and each feature has its own limits.
- Platform metrics. Every service has a Metrics tab with CPU, memory, disk usage and network traffic. It keeps up to 30 days of data, and dotted lines mark each deploy, so you can see which commit started a memory climb. With several replicas, the Sum view adds them together and the Replica view shows each one separately.
- Logs. Railway captures anything your process writes to stdout or stderr, plus build and deploy logs. Railway’s edge also writes an HTTP log for every request, with the path, status code and timings, and you can filter on all of them.
- Traces. For a traced service, Railway’s edge starts a trace on each request and passes a W3C
traceparentheader to your app. If the app runs an OpenTelemetry SDK, its spans join the same trace on the Traces page. Tracing is a preview feature. - Alerts. Project webhooks fire when a deploy changes state, including
FailedandCrashed, and when a monitor alerts. Monitors on the Observability dashboard alert when CPU, RAM, disk or network egress crosses a threshold. Monitors need the Pro plan.
What Railway leaves to you
For a Python API, these gaps matter. Railway does not collect application metrics such as request latency, error rate or business counts. Its docs send you to a third-party tool for those. The healthcheck only runs at the start of a deploy, and Railway never uses it to keep watching the service. There is also no log drain setting, so sending logs somewhere else takes a forwarder or an SDK. The sections below close each gap with what Railway already gives you, then show when an outside tool is worth adding.
Add a health route and bind Uvicorn to Railway’s port
Railway injects a PORT variable and uses it for health checks, so Uvicorn has to listen on 0.0.0.0 and $PORT. Add a cheap route that returns 200 once the app can serve requests:
# main.py
from fastapi import FastAPI
app = FastAPI()
@app.get("/health")
async def health():
return {"status": "ok"}
Then set the health check path and start command. You can do this in the service’s Settings, or in a railway.toml next to your code:
[deploy]
startCommand = "uvicorn main:app --host 0.0.0.0 --port $PORT --no-access-log"
healthcheckPath = "/health"
healthcheckTimeout = 300
restartPolicyType = "ON_FAILURE"
On each deploy, Railway calls /health until it gets a 2xx response, then moves traffic to the new version. If no 2xx arrives within the timeout, which is 300 seconds by default, Railway marks the deploy as failed. To give the app longer, for example when it loads a large model at startup, set the RAILWAY_HEALTHCHECK_TIMEOUT_SEC variable.
These details trip up FastAPI apps:
- The check comes from the hostname
healthcheck.railway.app. If you useTrustedHostMiddleware, add that host to the allowed list, or the check fails with a 400:
from fastapi.middleware.trustedhost import TrustedHostMiddleware
app.add_middleware(
TrustedHostMiddleware,
allowed_hosts=["api.example.com", "*.up.railway.app", "healthcheck.railway.app"],
)
- Keep
/healthshallow. If the route also queries Postgres, a short database blip turns into a failed deploy. Check dependencies from a separate route and watch that one from outside Railway (see the last setup section).
Railway has deprecated config-as-code files in favour of Infrastructure as Code. Existing railway.toml files keep working for legacy services until 2026-12-01. If you’d prefer not to keep a file, the same settings are in the dashboard.
Write one JSON log line per request to stdout
Railway treats a log line as structured when the whole JSON object sits on one line. It reads message and level, colours each line by level, and lets you filter on any other key as @key:value. Plain text lines are still stored. They get a level from their stream, info for stdout and error for stderr, so @level filters work on them, but they carry no custom attributes to filter on.
Python’s logging.StreamHandler writes to stderr unless you tell it otherwise, and Railway converts every stderr line to level: error. Point the handler at stdout and healthy requests stop showing up in red:
# logging_setup.py
import json
import logging
import sys
class JsonFormatter(logging.Formatter):
def format(self, record):
payload = {
"message": record.getMessage(),
"level": record.levelname.lower(),
"logger": record.name,
}
payload.update(getattr(record, "fields", {}))
if record.exc_info:
payload["exception"] = self.formatException(record.exc_info)
return json.dumps(payload, default=str)
def setup_logging():
handler = logging.StreamHandler(sys.stdout)
handler.setFormatter(JsonFormatter())
logging.basicConfig(level=logging.INFO, handlers=[handler], force=True)
The stack trace goes into a single exception field, so a crash stays as one entry you can open.
Next, log one event per request with the route, status, duration and trace ID:
# main.py
import logging
import time
from fastapi import FastAPI, Request
from opentelemetry import trace
from logging_setup import setup_logging
setup_logging()
log = logging.getLogger("app")
app = FastAPI()
@app.middleware("http")
async def log_requests(request: Request, call_next):
start = time.perf_counter()
fields = {"method": request.method, "path": request.url.path}
try:
response = await call_next(request)
except Exception:
fields["duration_ms"] = round((time.perf_counter() - start) * 1000, 1)
log.exception("unhandled error", extra={"fields": fields})
raise
ctx = trace.get_current_span().get_span_context()
fields.update(
status_code=response.status_code,
duration_ms=round((time.perf_counter() - start) * 1000, 1),
trace_id=format(ctx.trace_id, "032x") if ctx.is_valid else None,
)
log.info("request", extra={"fields": fields})
return response
trace_id stays empty until you turn on tracing in the next sections. After that, every log line links to its trace.
I built observability tools at Baselime and then led the Workers observability team at Cloudflare, and this is the habit I push hardest: write one wide event per request, with enough fields that you never need a second query to explain it.
The --no-access-log flag in the start command matters. Railway’s HTTP logs already record every request at the edge, and Uvicorn’s plain-text access lines only eat into the logging limit of 500 lines per second per replica. Above that limit, Railway drops lines and prints a warning with the number dropped.
Once the app is deployed, these queries work in the Log Explorer (the Observability tab) and in each deployment’s log panel:
@level:error
@path:/orders AND @status_code:>=500
@duration_ms:>500
@level:error AND "unhandled error"
In a terminal, railway logs shows the latest deployment’s logs.
Your plan sets how long Railway keeps logs, and traces are kept for the same period. If you upgrade, logs that had aged out come back straight away.
Railway log and trace retention by plan
| Plan | Retention |
|---|---|
| Free | 3 days |
| Trial | 7 days |
| Hobby | 7 days |
| Pro | 30 days |
| Enterprise | Up to 90 days |
Read latency and error rates from Railway’s HTTP logs
Before you write any metrics code, check the edge’s HTTP logs. They answer most latency and error questions. Each request has @httpStatus, @responseTime (time to first byte), @totalDuration, @path, @method and @srcIp, and the number fields accept comparisons and ranges:
@httpStatus:500..599
@path:/api/v1/orders AND @httpStatus:>=400
@responseTime:>1000
@totalDuration:>5000 @httpStatus:>=500
-@httpStatus:200
Railway only fills @responseDetails when the application fails to respond, so it explains requests that never reached a FastAPI handler. These filters give you error counts and slow requests per route, which covers most of what a small team asks. They can’t show where the time went inside the handler. Traces do that.
Turn on tracing and let OpenTelemetry instrument FastAPI
You turn tracing on per service and per environment. Once it’s on, the edge traces every request to the service, and Railway adds the OTEL_* variables on the next deploy. There is no sample rate: Railway traces every request.
- Enable tracing for the service in the Tracing setup panel on the Traces tab, or from the CLI:
railway trace enable --service api
railway trace status --all
- Install the OpenTelemetry distribution and let it detect FastAPI and your other libraries:
pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install
opentelemetry-bootstrap installs the matching packages, such as opentelemetry-instrumentation-fastapi. Run it locally, then add what it installed to requirements.txt or pyproject.toml so the Railway build installs them too.
- Put
opentelemetry-instrumentin front of the start command:
opentelemetry-instrument uvicorn main:app --host 0.0.0.0 --port $PORT --no-access-log
- Railway’s receiver only accepts traces, but the Python distribution exports metrics by default. Add these variables to the service:
OTEL_METRICS_EXPORTER=none
OTEL_LOGS_EXPORTER=none
- Redeploy, send a few requests, and wait for the App indicator in the Tracing setup panel to turn green.
Each request now produces a server span named after its method and route, with http.route and http.response.status_code. Calls made with httpx, requests or aiohttp become client spans. asyncpg, psycopg2, SQLAlchemy and redis produce spans that include the statement or command. To see your own slow work in the waterfall, wrap it in a span:
from opentelemetry import trace
tracer = trace.get_tracer("checkout")
def calculate_total(cart):
with tracer.start_as_current_span("calculate-total") as span:
span.set_attribute("cart.items", len(cart.items))
return price_items(cart)
The Traces page and the CLI use the same filter syntax, so you can find slow routes and errors from either one:
railway trace list --since 30m --errors
railway trace list --all --filter '@http.route:/checkout @duration:>500'
When a user reports a failed request, ask them for the x-railway-trace-id response header, or include it in your error responses. Paste it into the Trace ID field to open that exact request.
To skip the SDK, use the Automatic instrumentation switch. It traces Python processes with no code changes, on a best-effort basis. The SDK gives you custom spans and full control.
Tracing has limits. Each replica can export 1,000 spans per 10 seconds, and a search returns up to 500 traces. On a busy service, set OTEL_TRACES_SAMPLER=traceidratio and put a fraction in OTEL_TRACES_SAMPLER_ARG. The default parent-based sampler follows the edge, which traces everything.
Put the numbers on one dashboard and add monitors
- Open the Observability tab and check that the environment selector shows
production. The dashboard only covers one environment. - Click Start with a simple dashboard. Railway creates widgets for spend, service metrics and logs.
- Add a logs widget for the API service filtered to
@level:error, so errors sit next to the memory graph. - On the memory widget, open the three-dot menu, choose Add monitor, and set a threshold above your normal peak. Do the same for CPU.
- If the service has a volume, add a disk usage monitor. Add a network egress monitor if you watch bandwidth costs.
Monitors send email and in-app notifications. They also fire your project webhook, so they arrive in the same channel as failed deploys. For a Python service, the memory monitor matters most: a slow leak shows up on the graph well before the process crashes, and the alert gives you time to act.
Send failed and crashed deploys to Slack or Discord
You set webhooks per project, and they fire for every environment in that project.
- Create an incoming webhook in Slack (a
hooks.slack.comURL), or in a Discord channel under Integrations, then Webhooks. - In the Railway project, open Settings, then the Webhooks tab. Paste the URL, choose the events and click Save Webhook.
- Trigger a real deploy to check it works. Test Webhook sends from your browser, so CORS on the receiving end can make a working endpoint look broken.
Railway recognises Slack and Discord URLs and reformats the payload for them. Any other URL gets JSON with a type such as Deployment.failed and a resource block naming the project, environment and service. That’s enough for a small receiver to keep only production events. Host the receiver in its own project, so an outage in the app can’t take its alerting down with it.
By default, Railway restarts a crashed service on failure, up to 10 times, so a one-off crash often fixes itself before anyone looks. A failed deploy never fixes itself. Treat Deployment.failed as urgent, and page someone when the same service crashes repeatedly.
Cover the checks Railway stops running after deploy
The healthcheck stops once a deploy is live, so nothing on Railway tells you when the API stops answering overnight. For continuous checks, Railway’s docs point to the Uptime Kuma template in its marketplace. Run it in a separate project and point it at a route that checks your real dependencies.
For longer retention or application metrics, send telemetry to a third-party tool. Either path works:
- Vendor SDK. For Python, Datadog wraps the start command with
ddtrace-run, New Relic usesnewrelic-admin run-program, and you initialise Sentry in code. Map Railway’s variables to the vendor’s names, for exampleDD_SERVICE=${{RAILWAY_SERVICE_NAME}}andDD_VERSION=${{RAILWAY_DEPLOYMENT_ID}}. - OpenTelemetry to your own backend. Set
OTEL_EXPORTER_OTLP_ENDPOINTto the backend, and preferhttp/protobufon port 4318. If you set this variable yourself, Railway adds none of its tracing variables, so your app’s spans no longer reach the Traces page. The edge keeps tracing requests.
To forward raw stdout, run Vector or Fluent Bit as its own Railway service.
Where FastAPI monitoring on Railway goes wrong
- Uvicorn is not listening on the PORT variable. The health check returns service unavailable and the deploy fails. Start Uvicorn with
--host 0.0.0.0 --port $PORT. - Every log line is red. The handler writes to stderr. Pass
sys.stdouttoStreamHandler. - JSON logs show up as plain text. Pretty-printed or multi-line JSON breaks parsing. Write one object per line with
json.dumpsand no indent. - The health check fails with status 400.
TrustedHostMiddlewareis rejectinghealthcheck.railway.app. Add it toallowed_hosts. - Edge spans appear but the app’s spans don’t. The
OTEL_*variables arrive on the first deploy after you enable tracing, so redeploy. Also check that the service doesn’t set its ownOTEL_EXPORTER_OTLP_ENDPOINT. - The SDK logs export errors. It’s trying to send metrics or logs. Set
OTEL_METRICS_EXPORTER=noneandOTEL_LOGS_EXPORTER=none. - Staging shows nothing. Tracing, the dashboard and monitors are all per environment. Turn them on in each environment.
- Logs go missing under load. You’ve hit 500 lines per second per replica. Turn off Uvicorn access logs, raise the log level in production, and sample high-frequency events.
Go-live checklist for FastAPI on Railway
/healthreturns 200 and is set as the healthcheck path- Uvicorn listens on
0.0.0.0and$PORT - Logs are single-line JSON on stdout with
level,path,status_code,duration_msandtrace_id - Tracing is on in production and the App indicator is green
OTEL_METRICS_EXPORTERandOTEL_LOGS_EXPORTERare both set tonone- The production environment has a dashboard with an error log widget
- Memory and CPU monitors are set up (Pro plan)
- A project webhook posts to your alerts channel and has fired on a real deploy
- An outside uptime check hits a route that checks your dependencies
Let Polylane watch the Railway project for you
Polylane connects to Railway with OAuth or a workspace token. It syncs your projects, services and deployments, watches deploys, reads logs and checks service metrics for issues, with no alert rules to write or thresholds to tune. Railway’s logs and metrics are enough on their own as a telemetry source, and each confirmed issue goes to one agent that traces the cause and opens the fix as a pull request for you to review.
Running on Railway? See how Polylane monitors Railway in production.
Common questions.
Does Railway track request latency and error rates for a FastAPI app?
Not as metrics. The Metrics tab only covers CPU, memory, disk and network. Railway's HTTP logs do record `@httpStatus`, `@responseTime` and `@totalDuration` for every request, so filters like `@httpStatus:500..599` or `@responseTime:>1000` find errors and slow requests. For timing inside each route, enable tracing and run the app under `opentelemetry-instrument`.
Is Railway's healthcheck enough to know my API is up?
No. Railway only calls the healthcheck path at the start of a deploy, until it gets a 2xx, and never after the deploy is live. For continuous checks, Railway's docs suggest the Uptime Kuma template, or you can point any outside uptime service at your public domain.
How long does Railway keep logs, traces and metrics?
Railway keeps logs for 3 days on Free, 7 days on Trial and Hobby, 30 days on Pro and up to 90 days on Enterprise. Traces are kept for the same period. Metrics graphs hold up to 30 days of data. For longer history, send telemetry to a third-party tool with a vendor SDK or OpenTelemetry.
Why are all my FastAPI logs marked as errors on Railway?
By default, Python's logging StreamHandler writes to stderr, and Railway converts every stderr line to level error. Create the handler with `logging.StreamHandler(sys.stdout)` and include a `level` field in single-line JSON. Railway then colours each line by its real level.
Do I need the Pro plan to get alerts?
Monitors need the Pro plan. They are the threshold alerts on CPU, RAM, disk usage and network egress. Project webhooks cover failed builds, failed deploys and crashed deployments, and Railway formats them automatically for Slack and Discord URLs.
Can I send FastAPI traces to Datadog or Grafana and keep Railway's Traces page?
Not from the same exporter. If you set `OTEL_EXPORTER_OTLP_ENDPOINT` yourself, Railway adds none of its tracing variables, and your app's spans stop reaching the Traces page. The edge still traces requests. Choose one destination per service: Railway's built-in tracing, or a vendor setup such as `ddtrace-run` or an OTLP backend.
What does the Railway rate limit of 500 logs/sec warning mean?
Railway allows 500 log lines per second per replica on every plan and drops anything above that. For FastAPI, start Uvicorn with `--no-access-log`, because the edge HTTP logs already record every request. Then raise the production log level and sample high-frequency events.
How do I find the trace for one user's failed request?
Railway's edge adds an `x-railway-trace-id` header to every traced response. Have the client or your error page show it, then paste it into the Trace ID field on the Traces page, or run `railway trace get` with the ID.
Sources
- Railway Observability Dashboard
- Railway Logs
- Railway Metrics
- Railway Healthchecks
- Railway Tracing
- Railway Tracing for Python
- Set Up Alerts for Crashes, Restarts, and Failed Deploys
- Connect a Third-Party Observability Tool
- Railway Config as Code reference
- Product docs: Quickstart, Clouds and Detect
- Product overview and build log
Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.
More in this series
- 1 How to Monitor a Django App on Render
- 2 How to Monitor a FastAPI App on Railway: Logs, Traces and Alerts
- 3 How to Monitor a Supabase App in Production
- 4 How to Debug Cloudflare Workers Errors: Logs, Traces and Error Codes
- 5 How to debug Vercel function timeouts
- 6 Vercel 504 Gateway Timeout on Serverless Functions: Causes and Fixes
- 7 Cloudflare Workers error 1101: causes and how to fix it
- 8 Cloudflare Hyperdrive connection errors: causes and fixes
- 9 How to Monitor a Convex App in Production
Related
- What Is an AI SRE? How It Works and How to Evaluate One
An AI SRE is an AI agent that triages alerts, investigates incidents and proposes fixes. Learn how it works, the autonomy levels and how to evaluate one.
- Fix Railway “Application failed to respond” (502)
Why Railway returns “Application failed to respond” with a 502, how to tell a wrong host, port or target port from a crash or overload, and how to fix each.
- Fix Prisma “P1001: Can't reach database server”
Why Prisma prints “P1001: Can't reach database server”, how to test the host and port it names, and how to fix private hosts, IPv6, paused databases and URLs.
- How to monitor a vibe-coded app in production
Monitor a vibe-coded app in production: surface swallowed errors, add a health route, tag deploys, watch cron jobs and alert on what users feel first.
- Cloudflare Basin: a guide to the serverless data platform
Cloudflare Basin went GA on 1 October 2026. What its pipelines, Iceberg catalogue and SQL engine do, how to set it up, and the limits to plan for.