Get started Dashboard
Troubleshooting ·

Fix “error umounting /data: EBUSY” on Fly.io Machines

Explore with AI

error umounting /data: EBUSY: Device or resource busy, retrying in a bit

Fly's init process prints this line while it shuts a Machine down. It is trying to unmount the volume at /data while a process still has files open on it, and it retries a few times before the VM reboots. The line rarely needs fixing itself: find out what stopped the Machine (usually scale to zero on a Fly Postgres Development cluster, or Fly Proxy autostop) and turn that off, or give your process time to exit cleanly.

On this page

You usually find this line in fly logs right after Starting clean up. and Umounting /dev/vdb from /data, and just before reboot: Restarting system. It repeats a few times. Then the runner reports that the Machine exited and won’t be restarted. The app that used the volume, often a Fly Postgres database, now shows as stopped or suspended, and anything that connects to it starts failing.

The EBUSY is the last thing the Machine says before it goes down. It doesn’t explain why it went down. The fix is to find what started the shutdown, decide whether that should happen, and make sure the process that holds /data exits cleanly when it does.

What “error umounting /data: EBUSY” means

Every Fly Machine runs a small init process that Fly injects in front of your app. According to the Fly troubleshooting guide, init sets up networking and volume mounts, forwards signals to your app and coordinates clean shutdowns. You can’t disable or replace it.

When the Machine is told to stop, init runs a clean-up step. Part of that step unmounts the Fly Volume it mounted at boot. In the log, the volume’s block device is /dev/vdb and the mount path is /data, which is where Fly Postgres and most volume-backed apps keep their files.

EBUSY: Device or resource busy is the Linux kernel refusing the unmount, because something still has files open on that filesystem. Usually that’s your database or app process, still running or still writing. Init says retrying in a bit, tries again a few times, and then reboots the VM anyway.

So the line tells you three things:

  1. The Machine was shutting down on purpose. A crash would look different.
  2. Your main process hadn’t finished with /data when init tried to unmount it.
  3. Whether the Machine comes back depends on its restart policy and on whatever decided to stop it.

I built and led the Workers observability team at Cloudflare. On any platform that stops idle compute, the first question I ask of a shutdown log is what made the decision to stop. Read the lines before the EBUSY. They usually name the cause.

Find out what stopped the Machine

Pull the logs for the app that owns the volume, plus the Machine’s config:

fly logs -a my-app
fly status -a my-app
fly machine list -a my-app
fly m status -d <machine-id> -a my-app

fly m status -d prints the Machine config as JSON, including "restart": { "policy": ... }. With the log lines and that config, match your case to one of the causes below. They’re ordered by how often people hit them.

Cause 1: a Fly Postgres Development cluster scaled to zero

How to tell it’s yours: the app is an unmanaged Fly Postgres cluster created with the Development configuration. Just before Starting clean up., the log has a line reporting the current connection count. The shutdown happens on a regular cadence, and the database stays down until something connects.

The scale to zero docs explain the behaviour. New Development clusters can scale down after one hour. When the hour is up and there are no open connections, the database shuts down and waits until something tries to connect. When there are open connections, it stays up and checks again an hour later. The FLY_SCALE_TO_ZERO environment variable controls the feature, and the one hour timeout can’t be changed.

Check whether it’s on:

fly config save --app my-pg-app
grep FLY_SCALE_TO_ZERO fly.toml

Expected output when scale to zero is on:

  FLY_SCALE_TO_ZERO = "1h"

Fix: turn scale to zero off. These steps come from Fly’s docs:

  1. Save the config locally. The automated Postgres setup doesn’t write a fly.toml for you:
    fly config save --app my-pg-app
  2. Open fly.toml and delete this line from the [env] section:
    [env]
      FLY_SCALE_TO_ZERO = "1h"
  3. Check which image the cluster runs:
    fly image show --app my-pg-app
  4. Deploy with that same image. Fly’s example uses Postgres Flex 15.2. Use whatever tag step 3 showed:
    fly deploy . --image flyio/postgres-flex:15.2
  5. If the Machine is stopped right now, start it:
    fly machine start <machine-id> -a my-pg-app

If you’d rather keep scale to zero, the same docs say the apps that connect to the database may also need to scale to zero, because the database never sees zero connections while they hold one open. Those apps also need to wait while the database boots back up.

A warning before you invest more time here: Fly marks unmanaged Fly Postgres as archived. It’s no longer maintained, and Fly.io Support can’t help with it. For a database you depend on, plan a move to Fly.io Managed Postgres.

Cause 2: Fly Proxy autostop stopped an app that owns a volume

How to tell it’s yours: the app has an [http_service] or [[services]] block with auto_stop_machines = "stop" or "suspend", and the shutdowns follow quiet periods.

The autostop reference says Fly Proxy checks for excess capacity every few minutes. A region with one Machine and no traffic gets that Machine stopped or suspended. With several Machines, the proxy computes excess capacity = num of machines - (num machines over soft limit + 1) and stops one Machine when the result is 1 or more. The long-running tasks guide makes the key point: the proxy only counts traffic it can see. Background jobs, open database sessions on a private port, or a write in progress on /data don’t count. Your app also has no way to tell the proxy it’s busy.

grep -A6 '\[http_service\]' fly.toml

Fix: keep the volume-backed Machine running. Pick one:

  1. Turn autostop off for the app:
    [http_service]
      internal_port = 8080
      auto_stop_machines = "off"
      auto_start_machines = true
  2. Or keep autostop and keep one Machine warm, as the troubleshooting guide suggests:
    [http_service]
      internal_port = 8080
      auto_stop_machines = "stop"
      auto_start_machines = true
      min_machines_running = 1
  3. Or split the stateful work into its own process group with no [http_service] attached, so the proxy never stops it:
    [processes]
      web = "bundle exec puma"
      worker = "bundle exec sidekiq"
    
    [http_service]
      internal_port = 8080
      auto_stop_machines = "suspend"
      auto_start_machines = true
      processes = ["web"]
    Then scale each group separately with fly scale count web=2 worker=1.

Then deploy and check the result with fly config validate --strict. Without --strict, fly config validate quietly accepts keys it doesn’t recognise, so a typo can pass and then do nothing.

Suspend has its own limits. Machines must have 4 GB of RAM or less, a suspended Machine isn’t guaranteed to resume, and the clock can be wrong for a moment after resume. If the app checks timestamps at startup, use "stop".

Cause 3: your process doesn’t exit before init unmounts the volume

How to tell it’s yours: the shutdown has an expected trigger, such as fly deploy, fly machine stop or a host migration, but your app never logs its own shutdown message before Starting clean up.. Or it starts shutting down and the EBUSY lines come straight after.

According to the configuration reference, a controlled stop sends kill_signal to your main process, waits up to kill_timeout seconds, and then moves on to a forced shutdown. The default kill_timeout is 5 seconds and the maximum is 300. The reference names SIGINT as the default signal, while the long-running tasks guide lists SIGTERM. Set it explicitly so you don’t depend on either.

A database or app that can’t flush and close its files in 5 seconds is still holding /data when init tries the unmount.

Fix: give the process a signal it handles and time to use it.

  1. Set both keys at the top level of fly.toml:
    kill_signal = "SIGTERM"
    kill_timeout = "30s"
  2. Make sure the signal reaches your app. If the container starts through a shell wrapper such as CMD ["sh", "-c", "..."], the shell receives the signal and doesn’t pass it on. Use the exec form:
    CMD ["myapp"]
    Or end your entrypoint script with exec:
    #!/bin/sh
    ./run-migrations.sh
    exec myapp
  3. In your app, stop accepting new work when the signal arrives, finish or checkpoint what’s running, and exit a few seconds before kill_timeout runs out. The long-running tasks guide has drain examples for Node and Python.
  4. Deploy, run fly machine stop <machine-id>, and read the logs. Your app’s shutdown message should appear before Starting clean up., and the EBUSY retries should be gone.

Cause 4: the process exited and the restart policy left the Machine stopped

How to tell it’s yours: the log ends with the Machine exiting cleanly (exit code 0) and a note that it won’t restart. No autostop or scale to zero setting explains it.

The restart policy docs list three policies. on-fail, the default for Machines created by fly launch and fly deploy, restarts only after a non-zero exit, up to 10 retries in a 5-minute window. A clean exit means the Machine stays stopped. always restarts the Machine whatever the exit code, and Fly recommends it for always-on Machines with no services configured. no never restarts.

fly m status -d <machine-id> -a my-app

Fix: if the Machine should never stay down, set the policy:

fly m update <machine-id> --restart always -a my-app

Or set an app-wide default in fly.toml:

[[restart]]
  policy = "always"
  processes = ["app"]

Then find out why your process exited at all. A Machine that stops right after starting usually has no explicit CMD in its Dockerfile, or the app exits on a missing secret. Both show in fly logs.

Getting the stopped Machine running again

Whatever the cause, get service back first:

fly machine list -a my-app
fly machine start <machine-id> -a my-app

The Machine states reference lists stopped and suspended as persistent states. The Machine stays in them until you or the proxy start it. If it sits in starting or stopping for more than about 5 minutes, it may be stuck. The troubleshooting guide gives the order to try:

  1. fly machine restart <machine-id>
  2. fly machine update <machine-id> --yes --metadata foo=bar to nudge the platform state
  3. As a last resort, fly machine destroy --force <machine-id> followed by fly scale count <n>. Destroying a Machine is irreversible, so be sure the volume and snapshots are in order first.

Is the forced unmount a risk to the data on /data?

The reboot that follows the retries isn’t a clean unmount, so treat a repeating EBUSY as a reason to protect the volume. The troubleshooting guide says filesystem errors such as unable to read superblock mean the volume is corrupted, and that this can happen after a hard crash. Without snapshots, the data may be unrecoverable. Confirm you have snapshots and know the restore path:

fly volumes list -a my-app
fly volumes snapshots list <volume-id>
fly volumes create <name> --snapshot-id <snapshot-id> --region <region>

Keeping a volume-backed Machine from stopping unexpectedly

  • Decide on purpose whether the Machine may stop. For a database other apps depend on, turn off scale to zero and autostop. For an app that can stop, set min_machines_running and make the callers retry while it boots.
  • Size kill_timeout to your real shutdown time. Time a fly machine stop and give the setting a margin above it, up to the 300 second maximum.
  • Run fly config validate --strict in CI so a misspelt auto_stop_machines or kill_timeout fails the build.
  • Alert on the shutdown lines. Export your Fly logs to your log tool and alert on Starting clean up or error umounting for apps that should never stop. A database Machine shutting down at 3am deserves a page.
  • Keep volume snapshots on for any volume holding data you care about.

Catching a stopped Fly Machine with Polylane

Polylane connects Fly.io as a cloud account whose logs and metrics it can query, and watches Fly Machines alongside your other resources. When a check or a forwarded alert flags a stopped database, a fix run reads the logs around it, follows the connection to the apps that depend on it, and records the cause with its evidence. Code fixes arrive as a pull request for you to review.

Running on Fly.io? See how Polylane monitors Fly.io in production.

Common questions.

Is “error umounting /data: EBUSY” a bug in my app?

Not directly. Fly's init process prints it while shutting the Machine down, when a process still has files open on the volume mounted at /data. It's a sign that something stopped the Machine, often scale to zero or autostop, and that your process hadn't exited when init tried the unmount.

Why does my Fly Postgres app keep stopping by itself?

If it's an unmanaged Fly Postgres cluster created with the Development configuration, scale to zero is probably on. The database shuts down after one hour with no open connections, and that timeout can't be changed. Remove FLY_SCALE_TO_ZERO = "1h" from the [env] section of fly.toml and redeploy with the image that fly image show reports.

How do I start the stopped database Machine right now?

Run fly machine list -a my-pg-app to get the Machine ID, then fly machine start <machine-id> -a my-pg-app. If it stays in starting for more than about 5 minutes, try fly machine restart, then fly machine update <machine-id> --yes --metadata foo=bar.

What is the default kill_timeout on Fly.io, and how high can I set it?

The default is 5 seconds and the maximum is 300 seconds. Set it as a top-level key in fly.toml, for example kill_timeout = "30s", next to kill_signal = "SIGTERM". Fly treats it as best effort, so your app should still handle shorter stops.

Will setting auto_stop_machines to off cost more?

Yes. With autostop off, Machines stay up until something else stops them, so you pay for each one around the clock in every region you've scaled into. Fly bills stopped and suspended Machines the same way, but you don't pay for CPU and RAM while a Machine is stopped.

Why does my Machine stay stopped after a clean exit?

Machines created by fly launch and fly deploy default to the on-fail restart policy. It restarts only after a non-zero exit, up to 10 retries in a 5-minute window. To restart on any exit, run fly m update <machine-id> --restart always or add a [[restart]] block with policy = "always" to fly.toml.

Can the forced reboot after EBUSY corrupt my volume?

It can raise the risk, because the filesystem isn't unmounted cleanly. Fly's troubleshooting guide says hard crashes can corrupt a volume, showing errors such as unable to read superblock. Without snapshots, the data may be unrecoverable. Keep snapshots on and test a restore with fly volumes create --snapshot-id.

Should I still be running unmanaged Fly Postgres?

Fly marks unmanaged Fly Postgres as archived. It's no longer maintained, and Fly.io Support can't provide guidance for it. For production data, plan a move to Fly.io Managed Postgres.

Sources

  1. Scale to zero for Postgres Development projects (Fly Docs)
  2. Troubleshoot your deployment (Fly Docs)
  3. App configuration (fly.toml) (Fly Docs)
  4. Long-running tasks and machine lifecycle (Fly Docs)
  5. Fly Proxy autostop/autostart (Fly Docs)
  6. Machine restart policy (Fly Docs)
  7. Machine states and lifecycle (Fly Docs)
  8. fly machine (Fly Docs)
  9. Polylane documentation
  10. Polylane full content

About the author

Boris Tane

Founder of Polylane

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

Related

Nobody should be on-call. Polylane watches your infra, finds what broke, and writes the fix.

Get started for free