# Kubernetes CrashLoopBackOff: causes and fixes

> What CrashLoopBackOff means in Kubernetes, how to read the exit code and events behind it, and the fix for each common cause, from bad config to failing probes.

By Boris Tane, Founder of Polylane · Published October 5, 2026 · 11 min read
Canonical: https://polylane.com/learn/troubleshooting/kubernetes-crashloopbackoff/

```text
CrashLoopBackOff
```

CrashLoopBackOff means a container in the pod keeps exiting and the kubelet is waiting before it restarts it again, with a delay that starts at 10 seconds and doubles up to 5 minutes. It is a symptom: the cause is in the container's last exit. Run kubectl logs --previous and kubectl describe pod, read the Last State reason and exit code, then fix the app error, config, memory limit, probe or command it points to.

You run `kubectl get pods` and one row says `CrashLoopBackOff`. READY is `0/1`, the RESTARTS count keeps climbing, and the pod never serves traffic. If a rollout started it, the rollout stalls while the old pods keep serving.

CrashLoopBackOff is a waiting state. It tells you the container has exited several times and the kubelet is holding off before the next restart. It doesn't tell you why the container exits. That answer sits in the previous container's logs, its exit code and the pod's events, and this page shows you how to read each one.

## What CrashLoopBackOff means

The kubelet on the node prints it. When a container exits and the pod's `restartPolicy` allows a restart, the kubelet restarts it in the same pod on the same node. If it exits again, the kubelet applies an exponential back-off before each new attempt. The [Pod Lifecycle docs](https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#container-restarts) give the defaults: 10s, 20s, 40s and so on, capped at 300 seconds (5 minutes). Once a container has run for 10 minutes without problems, the kubelet resets the timer.

While the timer runs, the container sits in the Waiting state with the reason `CrashLoopBackOff`, and that reason is what kubectl shows in STATUS. The pod's phase stays `Running`. The docs call STATUS "a kubectl display field for user intuition" and warn you not to confuse it with the phase.

You'll see the same state in three places. In the kubelet source, [`doBackOff`](https://github.com/kubernetes/kubernetes/blob/master/pkg/kubelet/kuberuntime/kuberuntime_manager.go) records a Warning event with the reason `BackOff`:

```text
Warning  BackOff  12s (x8 over 2m)  kubelet  Back-off restarting failed container api in pod api-7d9c_default(3f1b...)
```

It also sets the container's waiting message to the current delay, which reads like this once the cap is reached:

```text
back-off 5m0s restarting failed container=api pod=api-7d9c_default(3f1b...)
```

And it returns the error `CrashLoopBackOff`, which is the reason you see in `kubectl describe pod` under `State: Waiting`.

```mermaid
graph TB
  A[Container exits] --> B[Kubelet restarts it]
  B --> C{Exits again?}
  C -->|yes| D[Wait 10s, doubling to 300s]
  D --> E[STATUS shows CrashLoopBackOff]
  E --> B
  C -->|runs 10 minutes| F[Back-off timer resets]
  style E fill:#fee2e2,stroke:#fca5a5,color:#7f1d1d
  style F fill:#d1fae5,stroke:#6ee7b7,color:#065f46
```

The delay explains why a crash loop feels slow to debug. After a few failures, each new attempt is up to 5 minutes away, so a fix you apply by hand can look like it did nothing for a while.

### Changing the back-off on newer clusters

Two kubelet features change these numbers. Neither fixes a crash, but they change how long you wait between attempts.

- **KubeletCrashLoopBackOffMax**, beta and enabled by default since v1.35. You set `maxContainerRestartPeriod` under `crashLoopBackOff` in the kubelet configuration, between `"1s"` and `"300s"`, per node. Delays still start at 10s and double, but stop at your cap. A cap under 10s becomes the starting delay too.
- **ReduceDefaultCrashLoopBackOffDecay**, alpha since v1.33 and off by default. With it on, restarts start at 1s and cap at 60s across the cluster. A node's own `maxContainerRestartPeriod` wins over these defaults.

```yaml
kind: KubeletConfiguration
crashLoopBackOff:
  maxContainerRestartPeriod: "100s"
```

## Reading the last exit before you change anything

Every cause below leaves a different trace. Collect three things first.

1. The logs of the container that crashed. Plain `kubectl logs` reads the current instance, which may be waiting and empty. `--previous` reads the one that died:
   ```bash
   kubectl logs <pod> -c <container> --previous
   ```
2. The container's last state and the pod's events:
   ```bash
   kubectl describe pod <pod>
   ```
   Look at `Last State: Terminated`, its `Reason` and `Exit Code`, then the Events table at the bottom.
3. The events on their own, if the describe output is long:
   ```bash
   kubectl events --for pod/<pod>
   kubectl get events --field-selector involvedObject.name=<pod>
   ```

The API server keeps events for one hour by default: the [kube-apiserver](https://kubernetes.io/docs/reference/command-line-tools-reference/kube-apiserver/) flag `--event-ttl` defaults to `1h0m0s`. The Last State in the container status survives longer, so read it even when the events are gone.

The exit code narrows the search fast. The [Bash manual](https://www.gnu.org/software/bash/manual/html_node/Exit-Status.html) says a process killed by signal N reports 128 plus N, a command that isn't found reports 127, and one that isn't executable reports 126. So 137 means SIGKILL (9) and 143 means SIGTERM (15).

```mermaid
graph TB
  A[kubectl describe pod] --> B{Last State reason?}
  B -->|OOMKilled, 137| C[Memory limit too low]
  B -->|Error, code 1| D[App exits: read previous logs]
  B -->|Error, 127 or 126| E[Command or entrypoint wrong]
  A --> F{Probe failed in Events?}
  F -->|yes| G[Liveness or startup probe]
  style C fill:#fee2e2,stroke:#fca5a5,color:#7f1d1d
  style G fill:#fee2e2,stroke:#fca5a5,color:#7f1d1d
```

## Why the container keeps exiting

The Pod Lifecycle docs list the usual causes: application errors that make the container exit, configuration errors such as wrong environment variables or missing files, resource constraints, and liveness or startup probes that fail. Here they are with the trace each one leaves, most common first.

### Cause 1: the application exits on startup

**How to tell it's yours:** `Last State` shows `Reason: Error` and a small non-zero exit code, often `1`. `kubectl logs --previous` ends with a stack trace, a panic or a message such as a failed database connection. The [Debug Running Pods guide](https://kubernetes.io/docs/tasks/debug/debug-application/debug-running-pod/#copying-a-pod-while-changing-its-command) shows exactly this shape:

```text
State:          Waiting
  Reason:       CrashLoopBackOff
Last State:     Terminated
  Reason:       Error
  Exit Code:    1
```

**Fix:**

1. Read the end of the previous logs. The last lines before the exit usually name the problem.
2. If the app exits before it logs anything, make Kubernetes keep its last output. With `terminationMessagePolicy: FallbackToLogsOnError`, the kubelet copies the last chunk of the log into the termination message when the container exits with an error, limited to 2048 bytes or 80 lines:
   ```yaml
   containers:
   - name: api
     image: registry.example.com/api:1.4.2
     terminationMessagePolicy: FallbackToLogsOnError
   ```
   Then read it from the pod status:
   ```bash
   kubectl get pod <pod> -o go-template='{{range .status.containerStatuses}}{{.lastState.terminated.message}}{{end}}'
   ```
3. If a deploy started the loop, roll back first and debug second:
   ```bash
   kubectl rollout history deployment/api
   kubectl rollout undo deployment/api
   kubectl rollout status deployment/api
   ```
   To go back to a specific revision, add `--to-revision=2`.
4. Reproduce it outside the cluster. The docs suggest running the same image locally, for example with `docker run --rm registry.example.com/api:1.4.2`, using the same environment variables the pod gets.

### Cause 2: bad configuration, environment variables or secrets

**How to tell it's yours:** the same image runs fine elsewhere, and the previous logs complain about a missing variable, an unreadable file or a value that doesn't parse. Often the crash started right after a change to a ConfigMap, a Secret or the Deployment's `env`.

**Fix:**

1. Compare what the pod gets with what the app expects:
   ```bash
   kubectl get pod <pod> -o jsonpath='{.spec.containers[0].env}'
   kubectl get pod <pod> -o jsonpath='{.spec.containers[0].envFrom}'
   kubectl describe configmap <name>
   ```
2. If the app has no shell to inspect, start a copy of the pod with the command replaced by one, as the debug guide shows. The `--container` flag replaces that container's command:
   ```bash
   kubectl debug <pod> -it --copy-to=<pod>-debug --container=api -- sh
   ```
   Inside, run `env`, check the mounted files and start the app by hand to watch it fail.
3. Fix the ConfigMap, Secret or `env` entry, then restart the Deployment so new pods pick it up:
   ```bash
   kubectl rollout restart deployment/api
   ```
4. Delete the debug copy when you're done: `kubectl delete pod <pod>-debug`.

### Cause 3: the container runs out of memory

**How to tell it's yours:** `Last State` shows `Reason: OOMKilled` and `Exit Code: 137`. The previous logs often stop mid-line, because the kernel killed the process without warning.

The container used more memory than its limit, and the kernel killed it. Each restart climbs back to the same usage and dies again, which turns into a crash loop. The fix is either a higher limit or a smaller footprint, and the [OOMKilled page](/learn/troubleshooting/kubernetes-oomkilled-exit-code-137/) covers both: reading `kubectl top pod --containers`, sizing requests and limits, and matching Node.js or JVM heap flags to the limit.

### Cause 4: a failing liveness or startup probe

**How to tell it's yours:** the logs look healthy, then the container stops. The Events show `Unhealthy` warnings followed by `Killing`. The [probes guide](https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/) shows the pattern:

```text
Warning  Unhealthy  10s (x3 over 20s)  kubelet  Liveness probe failed: cat: can't open '/tmp/healthy': No such file or directory
Normal   Killing    10s                kubelet  Container liveness failed liveness probe, will be restarted
```

When a liveness or startup probe fails, the kubelet kills the container and applies the restart policy. A slow-starting app with a tight liveness probe gets killed during startup, every time, and loops. Probe `timeoutSeconds` defaults to 1 second, which is short for an endpoint that touches a database.

**Fix:**

1. Check what the probe hits and whether the app answers in time:
   ```bash
   kubectl get pod <pod> -o jsonpath='{.spec.containers[0].livenessProbe}'
   ```
2. Add a startup probe that covers the slowest start you expect. Liveness and readiness probes don't run until it succeeds. The docs give this example, which allows up to 30 × 10 = 300 seconds to start:
   ```yaml
   ports:
   - name: liveness-port
     containerPort: 8080
   livenessProbe:
     httpGet:
       path: /healthz
       port: liveness-port
     failureThreshold: 1
     periodSeconds: 10
   startupProbe:
     httpGet:
       path: /healthz
       port: liveness-port
     failureThreshold: 30
     periodSeconds: 10
   ```
3. Keep the liveness endpoint cheap. It should answer whether this process can make progress. A check that calls a database or another service fails when that dependency is slow, and the kubelet restarts healthy containers as a result.
4. If the startup probe never succeeds, the container is killed after its window and the loop goes on. In that case the problem is the app, so go back to cause 1.

### Cause 5: the wrong command or entrypoint

**How to tell it's yours:** the container exits at once with no app logs at all. With a shell entrypoint, the exit code is `127` (command not found) or `126` (found but not executable). If the binary named in `command` doesn't exist in the image, the process never starts: read the Last State message and the Events for the container runtime's error.

Common triggers are a `command` that overrides the image's `ENTRYPOINT` with a wrong path, a script without execute permission, and a multi-stage build that left the binary out of the final image.

**Fix:**

1. Compare the pod's `command` and `args` with the image:
   ```bash
   kubectl get pod <pod> -o jsonpath='{.spec.containers[0].command} {.spec.containers[0].args}'
   docker image inspect registry.example.com/api:1.4.2 --format '{{.Config.Entrypoint}} {{.Config.Cmd}}'
   ```
2. Open the image with a shell and look for the file:
   ```bash
   kubectl debug <pod> -it --copy-to=<pod>-debug --container=api -- sh
   ls -l /app
   ```
3. For an image with no shell, attach an ephemeral debug container. `--target` points it at the process namespace of the named container, and the runtime has to support it:
   ```bash
   kubectl debug -it <pod> --image=busybox:1.28 --target=api
   ```
4. Fix the path or permissions in the manifest or Dockerfile (`RUN chmod +x /app/start.sh`), rebuild and redeploy.

## Confirming the crash loop is over

A single successful start proves little, because the container might crash again a minute later. Watch it:

```bash
kubectl get pod <pod> -w
```

Expected output once it's fixed:

```text
NAME        READY   STATUS    RESTARTS      AGE
api-7d9c    1/1     Running   6 (12m ago)   31m
```

Two signs matter. READY shows `1/1`, so the readiness probe passes. And the RESTARTS count stays flat for more than 10 minutes. That is the point where the kubelet resets the back-off timer, so the container has now run long enough to count as healthy. For a Deployment, `kubectl rollout status deployment/api` should report that the rollout finished.

## Stopping CrashLoopBackOff from coming back

- **Set a rollout deadline.** `.spec.progressDeadlineSeconds` defaults to 600. When a new version crash-loops, the Deployment gets a `Progressing` condition with reason `ProgressDeadlineExceeded`. Kubernetes only reports it, so alert on it or have your deploy tool roll back.
- **Give slow starters a startup probe** sized from a measured worst case, and keep liveness checks local to the process.
- **Validate config at startup and say what's missing.** An app that logs `DATABASE_URL is not set` before it exits turns a 20-minute hunt into a one-line fix.
- **Turn on `FallbackToLogsOnError`** for containers that can die before their logging is up.
- **Alert on restarts.** A container that restarts a few times an hour never shows CrashLoopBackOff for long, yet it is still failing. Alert on the restart count rising over a window, along with the status.
- **Ship events somewhere that keeps them.** With a one-hour default TTL, the evidence for an overnight crash loop is gone by morning.

The habit worth keeping is simple: look at the last exit, the last logs and the events together. With those three side by side, most crash loops take minutes.

## Tracing a crash loop back to its cause with Polylane

Polylane connects a cluster through an agent you install with Helm, with read-only RBAC limited to get, list and watch. Its agents inspect pods, deployments and nodes and read logs and events, so a crash loop is investigated from the same evidence this page uses, starting from the rollout that came before it.

Running on Kubernetes? See [how Polylane monitors Kubernetes in production](https://polylane.com/for/kubernetes/).

## Common questions

**How long does Kubernetes wait between restarts in CrashLoopBackOff?**

By default the kubelet waits 10 seconds after the first crash, then 20, then 40, doubling each time up to a cap of 300 seconds (5 minutes). Once a container has run for 10 minutes without problems, the kubelet resets the timer and the next crash counts as the first one.

**Can I shorten the CrashLoopBackOff delay?**

Yes, per node. The KubeletCrashLoopBackOffMax feature is beta and on by default since Kubernetes v1.35, so you can set crashLoopBackOff.maxContainerRestartPeriod in the kubelet configuration to anything from 1s to 300s. A separate alpha gate, ReduceDefaultCrashLoopBackOffDecay, changes the cluster defaults to a 1 second start and a 60 second cap, and it is off by default.

**Is CrashLoopBackOff a pod phase?**

No. It shows in the STATUS column of kubectl get pods and as the Waiting reason of a container, but the pod's phase is usually still Running. The Kubernetes docs warn against confusing that kubectl display field with the phase in the Pod API.

**Why does kubectl logs show nothing for my crashing pod?**

Plain kubectl logs reads the current container, which may be waiting in back-off and has printed nothing yet. Add --previous (or -p) to read the instance that crashed, and -c with the container name in a multi-container pod. If the app dies before it logs anything, set terminationMessagePolicy to FallbackToLogsOnError so the last 80 lines or 2048 bytes land in the pod status.

**What does exit code 137 mean in CrashLoopBackOff?**

137 is 128 plus 9, the number of SIGKILL. If the Last State reason is OOMKilled, the container went over its memory limit. If the reason is Error and the events show a failed liveness or startup probe, the kubelet killed it after the probe failed.

**My events are gone. How do I find out what happened?**

The API server keeps events for one hour by default (the kube-apiserver --event-ttl flag defaults to 1h0m0s). The container status still holds the last termination reason, exit code and message, so kubectl describe pod and kubectl get pod -o yaml still show Last State after the events expire.

**Should I just delete the pod?**

Deleting it makes the Deployment create a new pod with a fresh back-off timer, but the new container hits the same cause and loops again. Delete only after you've changed something: the image, the config, the limit or the probe.

**How do I roll back a deploy that started the crash loop?**

Run kubectl rollout history deployment/<name> to see the revisions, then kubectl rollout undo deployment/<name>, or add --to-revision=<n> for a specific one. Watch it with kubectl rollout status deployment/<name>.

## Sources

- [Pod Lifecycle (Kubernetes docs)](https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/)
- [kuberuntime_manager.go, doBackOff (Kubernetes source)](https://github.com/kubernetes/kubernetes/blob/master/pkg/kubelet/kuberuntime/kuberuntime_manager.go)
- [kubelet.go, crash loop back-off defaults (Kubernetes source)](https://github.com/kubernetes/kubernetes/blob/master/pkg/kubelet/kubelet.go)
- [Debug Running Pods (Kubernetes docs)](https://kubernetes.io/docs/tasks/debug/debug-application/debug-running-pod/)
- [Configure Liveness, Readiness and Startup Probes (Kubernetes docs)](https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/)
- [Determine the Reason for Pod Failure (Kubernetes docs)](https://kubernetes.io/docs/tasks/debug/debug-application/determine-reason-pod-failure/)
- [Deployments (Kubernetes docs)](https://kubernetes.io/docs/concepts/workloads/controllers/deployment/)
- [kubectl events (Kubernetes docs)](https://kubernetes.io/docs/reference/kubectl/generated/kubectl_events/)
- [kube-apiserver reference (Kubernetes docs)](https://kubernetes.io/docs/reference/command-line-tools-reference/kube-apiserver/)
- [kubectl rollout restart (Kubernetes docs)](https://kubernetes.io/docs/reference/kubectl/generated/kubectl_rollout/kubectl_rollout_restart/)
- [signal(7), Linux manual page](https://man7.org/linux/man-pages/man7/signal.7.html)
- [Exit Status (GNU Bash manual)](https://www.gnu.org/software/bash/manual/html_node/Exit-Status.html)
- [Kubernetes integration (Polylane docs)](https://docs.polylane.com/integrations/kubernetes)
- [Polylane documentation](https://docs.polylane.com/llms-full.txt)
- [Polylane full content](https://polylane.com/llms-full.txt)

## About the author

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

## Related

- [Kubernetes OOMKilled (exit code 137): causes and fixes](https://polylane.com/learn/troubleshooting/kubernetes-oomkilled-exit-code-137/): Why Kubernetes reports OOMKilled with exit code 137, how to tell a limit kill from a node eviction or a plain SIGKILL, and how to size memory so it stops.
- [Fix “context deadline exceeded” in Docker and fly deploy](https://polylane.com/learn/troubleshooting/docker-context-deadline-exceeded/): What “context deadline exceeded” means in Docker and fly deploy, how to find the call that timed out, and the fixes for WireGuard, builders and daemons.
- [Is It Safe to Give an AI SRE Read Access to Production?](https://polylane.com/learn/ai-in-production/is-it-safe-to-give-an-ai-sre-read-access-to-production/): Read access for an AI SRE is a sound first step if you scope it. The risks are secrets, customer data and exfiltration. Here are the controls and a checklist.
- [Fix “Durable Object reset because its code was updated”](https://polylane.com/learn/troubleshooting/cloudflare-durable-object-reset-because-its-code-was-updated/): Why a deploy makes Cloudflare Durable Objects throw “reset because its code was updated”, and how to retry safely, keep clients connected and lose no state.
- [Fix “remaining connection slots are reserved” in PostgreSQL](https://polylane.com/learn/troubleshooting/postgres-remaining-connection-slots-are-reserved/): Why Postgres refuses new logins with “remaining connection slots are reserved”, how to find who holds the connections, and how to pool and cap them for good.

Get started with one command: `curl -fsSL https://polylane.com/setup | bash` installs the CLI, connects your coding agents, and creates the account.
