Get started Dashboard
Reliability ·

Modal Clusters: multi-node GPUs and RDMA, step by step

Explore with AI

Modal Clusters became generally available on 1 October 2026. They let one Modal Function run as a gang-scheduled group of up to 32 full GPU nodes that start together and are billed by the second. To use one, put `@modal.clustered(size=N, rdma=True)` under an `@app.function` that requests every GPU on the node, read each container's rank and peer IPs from `modal.Cluster.from_context()`, and NCCL-based PyTorch gets RDMA at up to 6,400 Gbps per node on B300s or 3,200 Gbps on the other supported GPUs.

On this page

Modal announced on 1 October 2026 that Modal Clusters are generally available to every workspace. A single decorator, @modal.clustered, turns a Modal Function into a group of GPU containers. Modal places them together, starts them at the same moment and bills them by the second. Add rdma=True and the containers talk over an RDMA fabric at up to 6,400 Gbps per node on B300s, and Modal configures NCCL for you.

This guide explains what a cluster is and how it changes the way your code runs. It then walks through a first run with torchrun and with Ray, and covers the limits that Modal’s announcement and docs state: full nodes only, failures and preemptions that take down the whole cluster, plan GPU limits, and the libraries your image needs for RDMA.

What a Modal Cluster is

A Modal Cluster is a set of containers, one per node. Each container holds every GPU on its node, and Modal gang schedules the set: your code starts only after all the hardware you asked for has been acquired. The containers sit physically close together and share a private network. With RDMA on, they also share a high-bandwidth scale-out fabric (RoCE, InfiniBand or EFA, depending on the cloud underneath).

Modal runs your function in every container at once. Coordination is up to your code: decide who leads, find the peers, start the distributed framework. Modal’s docs say this job usually falls to SLURM or Kubernetes, and here it comes from a Python decorator.

The announcement says Modal tested the feature for 1.5 years before GA. It also needed a new scheduler. The normal Modal scheduler is greedy and decides node by node. The gang scheduler runs an observe, plan and act loop over the whole fleet. It collects every pending multi-node workload, assigns each one to nodes grouped by availability zone and network ID, adds capacity for anything that doesn’t fit, and notifies the chosen nodes. Clusters draw on the same capacity pool as everything else on Modal. That is why Modal says a cluster is running within seconds of the request.

How a clustered function behaves differently from a normal Modal function

Six behaviours change once you add @modal.clustered, and each one shapes how you write the code:

  • Every container gets the call’s arguments. Call a four-node Function once and your code runs four times in parallel.
  • Only rank 0’s return value reaches the caller. Return values from the other ranks are dropped. To combine results, send them to rank 0 over the private network or RDMA before it returns.
  • Clustered Servers send traffic only to rank 0. The leader passes requests on to the other containers.
  • The autoscaler adds and removes whole clusters. min_containers, max_containers and buffer_containers count individual nodes, and each must be a multiple of the cluster size. On a four-node cluster, min_containers=8 keeps two clusters warm.
  • Modal doesn’t keep the inputs in step across containers. Each container must avoid running ahead of the others. Before starting the next input, the leader should wait until every rank has finished the current one. The docs say this matters mostly for inference, and very little for a single training run.
  • One failure fails the whole call. If an input fails on any container, Modal stops the rest and marks the call as failed, even if other ranks succeeded.

The decorator also works on a class defined with @app.cls(), as long as the class has exactly one method. Web Functions are not supported. For HTTP workloads, use a clustered Server.

When a job needs more than one node

The announcement names two kinds of work. Training needs fast links between nodes for DDP all-reduce and for every kind of model parallelism. Inference with prefill-decode disaggregation needs to move the KV cache between nodes quickly.

Modal gives two examples of the data involved. Each training step of GLM 4.7 syncs about 717 GB of BF16 weights from the trainer to the rollout engine. That takes nearly two minutes over 50 Gbps TCP and under two seconds over RDMA, and it repeats for thousands of steps. Prefill-decode disaggregation on Llama 3.1 70B moves about 10 GB of KV cache for each 32k-token prompt, and it has to arrive within a time-to-first-token budget of a few hundred milliseconds.

The announcement names three early users. Decagon fine-tuned open models with up to a trillion parameters. 1x pre-trains the world model behind its NEO home robot on multi-node B300 clusters. Runway spreads video generation inference across several nodes. If your model and batch fit on a single eight-GPU node, a normal GPU Function is simpler and you don’t need any of this.

Starting a first cluster, step by step

  1. Install the SDK and sign in. Running GPUs needs a payment method on file.
pip install modal
python -m modal setup
modal token info
  1. Check your plan’s GPU limit before you pick a cluster size (see the sizing section below).

  2. Run a smoke test that only prints rank and addresses. It shows that gang scheduling and private networking work before you spend time on NCCL. It still bills two full H100 nodes while it runs, so keep it short.

import modal

app = modal.App("cluster-smoke-test")

@app.function(gpu="H100:8", timeout=10 * 60)
@modal.clustered(size=2)
def hello():
    cluster = modal.Cluster.from_context()
    rank = cluster.container_rank()
    ips = cluster.container_ips()
    role = "(main)" if rank == 0 else ""
    print(f"container_rank={rank} {role} world_size={len(ips)} main_addr={ips[0]!r}")
    return rank

@app.local_entrypoint()
def main():
    print("returned:", hello.remote())
modal run smoke_test.py

You should see one log line from each container and a single return value, which comes from rank 0:

container_rank=0 (main) world_size=2 main_addr='fdaa:...'
container_rank=1  world_size=2 main_addr='fdaa:...'
returned: 0

The addresses are IPv6 by default. Each workspace gets its own private subnet under the fdaa::/16 prefix.

  1. Build an image that can use RDMA. Start from an official CUDA base image and install the InfiniBand verbs library. This is the pattern from Modal’s docs:
cuda_version = "12.9.1"
flavor = "devel"
operating_sys = "ubuntu22.04"
tag = f"{cuda_version}-{flavor}-{operating_sys}"

image = (
    modal.Image.from_registry(f"nvidia/cuda:{tag}", add_python="3.12")
    .apt_install("libibverbs1")
    .uv_pip_install("torch")
    .add_local_dir("train", remote_path="/root/train")
)
  1. Turn on RDMA and start your framework. The torchrun and Ray patterns follow.

  2. Set a timeout, a retry policy and checkpoints. The default timeout is 300 seconds, which is far too short for training (see the failure section below).

  3. Measure the fabric once. Modal publishes an RDMA performance test in the modal-labs/multinode-training-guide repository, under benchmark. Run it before you blame slow steps on your model code.

torchrun across two H100 nodes

N_NODES = 2
N_GPUS_PER_NODE = 8

@app.function(gpu=f"H100:{N_GPUS_PER_NODE}", image=image, timeout=60 * 60)
@modal.clustered(size=N_NODES, rdma=True)
def run():
    from torch.distributed.run import parse_args, run

    cluster = modal.Cluster.from_context()
    args = [
        f"--nnodes={N_NODES}",
        f"--nproc-per-node={N_GPUS_PER_NODE}",
        f"--node-rank={cluster.container_rank()}",
        f"--master-addr={cluster.container_ips()[0]}",
        "/root/train/benchmark.py",
    ]
    run(parse_args(args))

Rank 0’s address becomes the torchrun master address, and every container passes its own rank as the node rank.

Ray needs IPv4 addresses

Ray only works with IPv4, so ask for IPv4 addresses with container_ips(family="ipv4"). Start the Ray head on rank 0 and a worker on every other rank:

@app.function(gpu="H100:8", image=image, timeout=24 * 60 * 60)
@modal.clustered(size=2, rdma=True)
async def train():
    cluster = modal.Cluster.from_context()
    ips = cluster.container_ips(family="ipv4")
    rank = cluster.container_rank()
    os.environ["HOST_IP"] = ips[rank]
    if rank == 0:
        await run_head(ips[rank])
    else:
        await run_worker(ips[rank], ips[0])

In Modal’s example, the head starts ray start --head and retries ray.init(address="auto") until it connects. It then submits the job through JobSubmissionClient and opens the dashboard on port 8265 with modal.forward(8265), streaming job logs to the Modal dashboard. Workers run ray start --address against the head on port 6379, then loop to stay alive while the job runs.

What rdma=True sets up and what your image still carries

6.4 Tbps
RDMA bandwidth per node on B300 clusters
Other RDMA-capable GPUs get 3,200 Gbps. Plain i6pn private networking between containers runs at 50 Gbps or more. Source: Modal Multi-node Clusters and Cluster networking docs.

RDMA lets one host read and write another host’s GPU or CPU memory over the network without going through the kernel. Every cloud does it a little differently, with its own drivers, environment variables and userspace libraries. The point of rdma=True is that you don’t have to deal with that.

  • Modal sets up the hosts. It installs the cloud’s drivers and configures the NIC that sits next to each GPU. If you use NCCL or a framework built on it, Modal also sets the NCCL environment variables in your containers.
  • Your image carries the userspace libraries. That usually means libcudart.so, libibverbs.so.1 and libmlx5.so.1. The simplest route is a CUDA base image plus apt_install("libibverbs1").
  • Only some GPUs support it. RDMA works on H100, H200, B200, B300 and GB200.
  • The interface depends on the hardware. You get either InfiniBand Verbs or EFA. If your workload can’t use EFA, the experimental option experimental_options={"efa_disabled": True} forces an InfiniBand Verbs host. Modal warns that this may lengthen scheduling and reduce available capacity.
  • It runs inside gVisor. Modal runs containers in the gVisor sandbox. When NCCL registers memory, gVisor intercepts the call at the sandbox boundary and hands it to the host kernel, which pins GPU memory and gives the NIC its address. Modal built this RDMA support into gVisor and upstreamed it.

Sizing a cluster against plan limits and the per-second bill

The docs allow up to 32 nodes and 256 GPUs per cluster, and say to contact Modal for anything larger. The announcement adds that your plan’s GPU limits cap cluster size. Here is how the two self-serve plans compare:

GPU concurrency

GPUs
Starter plan 10
Team plan 50

Containers

containers
Starter plan 100
Team plan 5,000
Figure 1
GPU and container concurrency on Modal's Starter and Team plans

Cluster nodes must use every GPU on the node, so even the smallest cluster of two eight-GPU nodes needs more GPUs than the Starter plan’s 10. The Team plan costs $250 a month plus compute and allows 50. Enterprise lists higher, custom GPU concurrency.

Compute is billed per GPU-second. The pricing page lists the Nvidia B300 at $0.001972 per second and the H100 SXM5 at $0.001097 per second. Choosing a specific region multiplies base prices by 1.15 to 1.75x. Modal’s announcement contrasts this with other on-demand cluster offers, which charge by the hour or require a reservation.

Failures, preemption and timeouts hit the whole cluster

A long multi-node job must survive two Modal behaviours that apply to the cluster as a unit:

  • A preemption hits every node. Modal stops all containers in the cluster and retries the call with the same input. Preemptions are rare, but the chance of one grows the longer a job runs. Setting nonpreemptible doesn’t help here, because it isn’t supported for GPU Functions.
  • A failure on one rank fails the whole call. Set a retry policy so Modal reruns the call for you.

Both mean every restart begins from zero unless you checkpoint. Write checkpoints to a Volume at regular intervals, make each step safe to repeat, and use the exit handler (Modal sends an interrupt signal on preemption) to save state within the grace period.

@app.function(
    gpu="B300:8",
    image=image,
    timeout=60 * 60 * 24,
    startup_timeout=30 * 60,
    retries=modal.Retries(initial_delay=0.0, max_retries=10),
)
@modal.clustered(size=4, rdma=True)
def train_model():
    ...

Timeouts measure execution time, from 1 second up to 24 hours, and each retry starts a fresh timeout. When it runs out, the caller gets modal.exception.FunctionTimeoutError. startup_timeout covers container startup separately, which helps when loading weights or importing packages is slow. Before client v1.1.4, timeout covered both.

Reading GPU and RDMA health from container logs

Modal monitors GPU health on every host and, according to the announcement, now covers RDMA health too. Problems show up in your container logs in this format:

[gpu-health] [LEVEL] GPU-[UUID]: EVENT_TYPE: MSG
[gpu-health] [CRITICAL] GPU-1234: XID: NVRM: Xid (PCI:0000:c6:00): 79, pid=1101234, name=nvc:[driver], GPU has fallen off the bus.

On a CRITICAL event, Modal drains the worker and moves your containers. On a WARN event, Modal does nothing. A warning may be harmless, or it may point to a bug in your application or a library. SXid events report faults in NVSwitch, the GPU-to-GPU interconnect, so they only matter in multi-GPU containers, which covers every cluster node. When your code catches a GPU fault, stop the container from taking new inputs:

import modal.experimental

try:
    ...  # code that may hit a GPU fault, such as an illegal memory access
except RuntimeError:
    modal.experimental.stop_fetching_inputs()
    return

I built and led the Workers observability team at Cloudflare, and the habit I would carry over is simple: send these log lines to the tool you already alert from before the first long run. Modal documents integrations for Datadog and for OpenTelemetry providers. A WARN that nobody reads at hour three can easily become the restart that costs you hour twenty.

Gotchas that stop a cluster from starting or finishing

  • Partial nodes are rejected. H100:4 is invalid for a cluster and H100:8 is valid. CPU-only clustered functions aren’t supported. A10 nodes top out at four GPUs, but A10s don’t support RDMA anyway.
  • B300 needs CUDA 13.1 or later. The RDMA image example in Modal’s docs pins CUDA 12.9.1. For B300 nodes, pick a CUDA 13 base image and libraries built for it. The same applies to gpu="B200+", which may land on a B300.
  • H100 requests can land on H200s. Modal may upgrade gpu="H100" to an H200 at no extra cost. Use H100! when you benchmark and need the hardware fixed.
  • Ray fails on IPv6. container_ips() returns IPv6 unless you pass family="ipv4".
  • Results from other ranks disappear. Only rank 0’s return value comes back. Gather results on rank 0 first.
  • Autoscaling settings that aren’t multiples of the cluster size are invalid, because those settings count nodes.
  • The cluster can’t be reached from outside. i6pn addresses are private to the workspace and limited to one region. To expose a dashboard or endpoint, open a modal.Tunnel or modal.forward from one container, usually rank 0.
  • The decorator name differs between pages. The GA announcement and the Clusters docs use @modal.clustered. The cluster networking page still refers to the earlier Beta name, @modal.experimental.clustered. The announcement’s snippet also calls cluster.private_ips(), while the Clusters docs document container_ips(). Write new code against the Clusters docs.

A checklist before your first multi-node run

  • The plan’s GPU concurrency covers nodes × 8, or you’ve talked to Modal about a larger limit.
  • The GPU type supports RDMA (H100, H200, B200, B300 or GB200), and the image’s CUDA version matches it.
  • The image includes libibverbs1 on a CUDA base image, and the RDMA benchmark has run once.
  • The leader address and node rank come from modal.Cluster.from_context(), and Ray uses IPv4.
  • timeout, startup_timeout and modal.Retries are set, and checkpoints go to a Volume.
  • Rank 0 gathers results, and for inference it waits for every rank before starting the next input.
  • [gpu-health] log lines go to a tool someone actually watches.

Running on Modal? See how Polylane monitors Modal in production.

Common questions.

When did Modal Clusters become generally available?

Modal announced general availability on 1 October 2026, in a post by Peyton Walters on the Modal blog. Clusters are open to every workspace, and Modal says it tested them for 1.5 years before GA. You enable them with the `@modal.clustered` decorator.

Which GPUs support RDMA on Modal Clusters?

RDMA works on H100, H200, B200, B300 and GB200 GPUs. B300 clusters get 6,400 Gbps of RDMA networking per node and the others get 3,200 Gbps. Clusters without RDMA still get Modal's private i6pn network, which runs at 50 Gbps or more.

How large can a Modal Cluster be?

The docs allow up to 32 nodes and 256 GPUs per cluster, and say to contact Modal for anything larger. Your plan's GPU limits also cap cluster size. The Starter plan allows 10 concurrent GPUs, Team allows 50, and Enterprise is custom.

Can a cluster node use fewer than eight GPUs?

No. Clustered functions must request every GPU on the node, so `H100:8` is valid and `H100:4` is rejected. CPU-only clustered functions aren't supported either.

What happens if one node fails or is preempted?

If an input fails on any container, Modal stops all the others and marks the whole call as failed. A preemption also stops the entire cluster, and Modal retries it with the same input. Set `modal.Retries` and write checkpoints to a Volume, because the `nonpreemptible` option isn't supported for GPU Functions.

Do I have to configure NCCL or InfiniBand myself?

No. With `rdma=True`, Modal installs the cloud's drivers, configures the NICs and sets the NCCL environment variables. Your image still needs the userspace libraries, usually `libcudart.so`, `libibverbs.so.1` and `libmlx5.so.1`. A CUDA base image plus `apt_install("libibverbs1")` covers them.

How are Modal Clusters billed?

By the second, at the normal per-GPU rates, with no hourly minimum or reservation. The pricing page lists the B300 at $0.001972 per second and the H100 SXM5 at $0.001097 per second. Choosing a specific region multiplies base prices by 1.15 to 1.75x.

Can I serve HTTP inference from a cluster?

Yes, through a clustered Server. Web Functions aren't supported with `@modal.clustered`. A clustered Server receives traffic only on rank 0, which must pass requests to the other containers and keep them in step.

Sources

  1. Modal Clusters are generally available (Modal blog)
  2. Multi-node Clusters (Modal Docs)
  3. Multi-node training (Modal Docs)
  4. Cluster networking (Modal Docs)
  5. GPU acceleration (Modal Docs)
  6. Preemption (Modal Docs)
  7. Timeouts (Modal Docs)
  8. GPU Health (Modal Docs)
  9. Getting started with Modal (Modal Docs)
  10. Modal pricing

About the author

Boris Tane

Founder of Polylane

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

Related

Nobody should be on-call. Polylane watches your infra, finds what broke, and writes the fix.

Get started for free