Modal Clusters: multi-node GPUs and RDMA, step by step
Explore with AI
Modal Clusters became generally available on 1 October 2026. They let one Modal Function run as a gang-scheduled group of up to 32 full GPU nodes that start together and are billed by the second. To use one, put `@modal.clustered(size=N, rdma=True)` under an `@app.function` that requests every GPU on the node, read each container's rank and peer IPs from `modal.Cluster.from_context()`, and NCCL-based PyTorch gets RDMA at up to 6,400 Gbps per node on B300s or 3,200 Gbps on the other supported GPUs.
On this page
- What a Modal Cluster is
- How a clustered function behaves differently from a normal Modal function
- When a job needs more than one node
- Starting a first cluster, step by step
- What rdma=True sets up and what your image still carries
- Sizing a cluster against plan limits and the per-second bill
- Failures, preemption and timeouts hit the whole cluster
- Reading GPU and RDMA health from container logs
- Gotchas that stop a cluster from starting or finishing
- A checklist before your first multi-node run
- Common questions
Modal announced on 1 October 2026 that Modal Clusters are generally available to every workspace. A single decorator, @modal.clustered, turns a Modal Function into a group of GPU containers. Modal places them together, starts them at the same moment and bills them by the second. Add rdma=True and the containers talk over an RDMA fabric at up to 6,400 Gbps per node on B300s, and Modal configures NCCL for you.
This guide explains what a cluster is and how it changes the way your code runs. It then walks through a first run with torchrun and with Ray, and covers the limits that Modal’s announcement and docs state: full nodes only, failures and preemptions that take down the whole cluster, plan GPU limits, and the libraries your image needs for RDMA.
What a Modal Cluster is
A Modal Cluster is a set of containers, one per node. Each container holds every GPU on its node, and Modal gang schedules the set: your code starts only after all the hardware you asked for has been acquired. The containers sit physically close together and share a private network. With RDMA on, they also share a high-bandwidth scale-out fabric (RoCE, InfiniBand or EFA, depending on the cloud underneath).
Modal runs your function in every container at once. Coordination is up to your code: decide who leads, find the peers, start the distributed framework. Modal’s docs say this job usually falls to SLURM or Kubernetes, and here it comes from a Python decorator.
The announcement says Modal tested the feature for 1.5 years before GA. It also needed a new scheduler. The normal Modal scheduler is greedy and decides node by node. The gang scheduler runs an observe, plan and act loop over the whole fleet. It collects every pending multi-node workload, assigns each one to nodes grouped by availability zone and network ID, adds capacity for anything that doesn’t fit, and notifies the chosen nodes. Clusters draw on the same capacity pool as everything else on Modal. That is why Modal says a cluster is running within seconds of the request.
How a clustered function behaves differently from a normal Modal function
Six behaviours change once you add @modal.clustered, and each one shapes how you write the code:
- Every container gets the call’s arguments. Call a four-node Function once and your code runs four times in parallel.
- Only rank 0’s return value reaches the caller. Return values from the other ranks are dropped. To combine results, send them to rank 0 over the private network or RDMA before it returns.
- Clustered Servers send traffic only to rank 0. The leader passes requests on to the other containers.
- The autoscaler adds and removes whole clusters.
min_containers,max_containersandbuffer_containerscount individual nodes, and each must be a multiple of the cluster size. On a four-node cluster,min_containers=8keeps two clusters warm. - Modal doesn’t keep the inputs in step across containers. Each container must avoid running ahead of the others. Before starting the next input, the leader should wait until every rank has finished the current one. The docs say this matters mostly for inference, and very little for a single training run.
- One failure fails the whole call. If an input fails on any container, Modal stops the rest and marks the call as failed, even if other ranks succeeded.
The decorator also works on a class defined with @app.cls(), as long as the class has exactly one method. Web Functions are not supported. For HTTP workloads, use a clustered Server.
When a job needs more than one node
The announcement names two kinds of work. Training needs fast links between nodes for DDP all-reduce and for every kind of model parallelism. Inference with prefill-decode disaggregation needs to move the KV cache between nodes quickly.
Modal gives two examples of the data involved. Each training step of GLM 4.7 syncs about 717 GB of BF16 weights from the trainer to the rollout engine. That takes nearly two minutes over 50 Gbps TCP and under two seconds over RDMA, and it repeats for thousands of steps. Prefill-decode disaggregation on Llama 3.1 70B moves about 10 GB of KV cache for each 32k-token prompt, and it has to arrive within a time-to-first-token budget of a few hundred milliseconds.
The announcement names three early users. Decagon fine-tuned open models with up to a trillion parameters. 1x pre-trains the world model behind its NEO home robot on multi-node B300 clusters. Runway spreads video generation inference across several nodes. If your model and batch fit on a single eight-GPU node, a normal GPU Function is simpler and you don’t need any of this.
Starting a first cluster, step by step
- Install the SDK and sign in. Running GPUs needs a payment method on file.
pip install modal
python -m modal setup
modal token info
-
Check your plan’s GPU limit before you pick a cluster size (see the sizing section below).
-
Run a smoke test that only prints rank and addresses. It shows that gang scheduling and private networking work before you spend time on NCCL. It still bills two full H100 nodes while it runs, so keep it short.
import modal
app = modal.App("cluster-smoke-test")
@app.function(gpu="H100:8", timeout=10 * 60)
@modal.clustered(size=2)
def hello():
cluster = modal.Cluster.from_context()
rank = cluster.container_rank()
ips = cluster.container_ips()
role = "(main)" if rank == 0 else ""
print(f"container_rank={rank} {role} world_size={len(ips)} main_addr={ips[0]!r}")
return rank
@app.local_entrypoint()
def main():
print("returned:", hello.remote())
modal run smoke_test.py
You should see one log line from each container and a single return value, which comes from rank 0:
container_rank=0 (main) world_size=2 main_addr='fdaa:...'
container_rank=1 world_size=2 main_addr='fdaa:...'
returned: 0
The addresses are IPv6 by default. Each workspace gets its own private subnet under the fdaa::/16 prefix.
- Build an image that can use RDMA. Start from an official CUDA base image and install the InfiniBand verbs library. This is the pattern from Modal’s docs:
cuda_version = "12.9.1"
flavor = "devel"
operating_sys = "ubuntu22.04"
tag = f"{cuda_version}-{flavor}-{operating_sys}"
image = (
modal.Image.from_registry(f"nvidia/cuda:{tag}", add_python="3.12")
.apt_install("libibverbs1")
.uv_pip_install("torch")
.add_local_dir("train", remote_path="/root/train")
)
-
Turn on RDMA and start your framework. The torchrun and Ray patterns follow.
-
Set a timeout, a retry policy and checkpoints. The default timeout is 300 seconds, which is far too short for training (see the failure section below).
-
Measure the fabric once. Modal publishes an RDMA performance test in the
modal-labs/multinode-training-guiderepository, underbenchmark. Run it before you blame slow steps on your model code.
torchrun across two H100 nodes
N_NODES = 2
N_GPUS_PER_NODE = 8
@app.function(gpu=f"H100:{N_GPUS_PER_NODE}", image=image, timeout=60 * 60)
@modal.clustered(size=N_NODES, rdma=True)
def run():
from torch.distributed.run import parse_args, run
cluster = modal.Cluster.from_context()
args = [
f"--nnodes={N_NODES}",
f"--nproc-per-node={N_GPUS_PER_NODE}",
f"--node-rank={cluster.container_rank()}",
f"--master-addr={cluster.container_ips()[0]}",
"/root/train/benchmark.py",
]
run(parse_args(args))
Rank 0’s address becomes the torchrun master address, and every container passes its own rank as the node rank.
Ray needs IPv4 addresses
Ray only works with IPv4, so ask for IPv4 addresses with container_ips(family="ipv4"). Start the Ray head on rank 0 and a worker on every other rank:
@app.function(gpu="H100:8", image=image, timeout=24 * 60 * 60)
@modal.clustered(size=2, rdma=True)
async def train():
cluster = modal.Cluster.from_context()
ips = cluster.container_ips(family="ipv4")
rank = cluster.container_rank()
os.environ["HOST_IP"] = ips[rank]
if rank == 0:
await run_head(ips[rank])
else:
await run_worker(ips[rank], ips[0])
In Modal’s example, the head starts ray start --head and retries ray.init(address="auto") until it connects. It then submits the job through JobSubmissionClient and opens the dashboard on port 8265 with modal.forward(8265), streaming job logs to the Modal dashboard. Workers run ray start --address against the head on port 6379, then loop to stay alive while the job runs.
What rdma=True sets up and what your image still carries
RDMA lets one host read and write another host’s GPU or CPU memory over the network without going through the kernel. Every cloud does it a little differently, with its own drivers, environment variables and userspace libraries. The point of rdma=True is that you don’t have to deal with that.
- Modal sets up the hosts. It installs the cloud’s drivers and configures the NIC that sits next to each GPU. If you use NCCL or a framework built on it, Modal also sets the NCCL environment variables in your containers.
- Your image carries the userspace libraries. That usually means
libcudart.so,libibverbs.so.1andlibmlx5.so.1. The simplest route is a CUDA base image plusapt_install("libibverbs1"). - Only some GPUs support it. RDMA works on H100, H200, B200, B300 and GB200.
- The interface depends on the hardware. You get either InfiniBand Verbs or EFA. If your workload can’t use EFA, the experimental option
experimental_options={"efa_disabled": True}forces an InfiniBand Verbs host. Modal warns that this may lengthen scheduling and reduce available capacity. - It runs inside gVisor. Modal runs containers in the gVisor sandbox. When NCCL registers memory, gVisor intercepts the call at the sandbox boundary and hands it to the host kernel, which pins GPU memory and gives the NIC its address. Modal built this RDMA support into gVisor and upstreamed it.
Sizing a cluster against plan limits and the per-second bill
The docs allow up to 32 nodes and 256 GPUs per cluster, and say to contact Modal for anything larger. The announcement adds that your plan’s GPU limits cap cluster size. Here is how the two self-serve plans compare:
GPU concurrency
Containers
Cluster nodes must use every GPU on the node, so even the smallest cluster of two eight-GPU nodes needs more GPUs than the Starter plan’s 10. The Team plan costs $250 a month plus compute and allows 50. Enterprise lists higher, custom GPU concurrency.
Compute is billed per GPU-second. The pricing page lists the Nvidia B300 at $0.001972 per second and the H100 SXM5 at $0.001097 per second. Choosing a specific region multiplies base prices by 1.15 to 1.75x. Modal’s announcement contrasts this with other on-demand cluster offers, which charge by the hour or require a reservation.
Failures, preemption and timeouts hit the whole cluster
A long multi-node job must survive two Modal behaviours that apply to the cluster as a unit:
- A preemption hits every node. Modal stops all containers in the cluster and retries the call with the same input. Preemptions are rare, but the chance of one grows the longer a job runs. Setting
nonpreemptibledoesn’t help here, because it isn’t supported for GPU Functions. - A failure on one rank fails the whole call. Set a retry policy so Modal reruns the call for you.
Both mean every restart begins from zero unless you checkpoint. Write checkpoints to a Volume at regular intervals, make each step safe to repeat, and use the exit handler (Modal sends an interrupt signal on preemption) to save state within the grace period.
@app.function(
gpu="B300:8",
image=image,
timeout=60 * 60 * 24,
startup_timeout=30 * 60,
retries=modal.Retries(initial_delay=0.0, max_retries=10),
)
@modal.clustered(size=4, rdma=True)
def train_model():
...
Timeouts measure execution time, from 1 second up to 24 hours, and each retry starts a fresh timeout. When it runs out, the caller gets modal.exception.FunctionTimeoutError. startup_timeout covers container startup separately, which helps when loading weights or importing packages is slow. Before client v1.1.4, timeout covered both.
Reading GPU and RDMA health from container logs
Modal monitors GPU health on every host and, according to the announcement, now covers RDMA health too. Problems show up in your container logs in this format:
[gpu-health] [LEVEL] GPU-[UUID]: EVENT_TYPE: MSG
[gpu-health] [CRITICAL] GPU-1234: XID: NVRM: Xid (PCI:0000:c6:00): 79, pid=1101234, name=nvc:[driver], GPU has fallen off the bus.
On a CRITICAL event, Modal drains the worker and moves your containers. On a WARN event, Modal does nothing. A warning may be harmless, or it may point to a bug in your application or a library. SXid events report faults in NVSwitch, the GPU-to-GPU interconnect, so they only matter in multi-GPU containers, which covers every cluster node. When your code catches a GPU fault, stop the container from taking new inputs:
import modal.experimental
try:
... # code that may hit a GPU fault, such as an illegal memory access
except RuntimeError:
modal.experimental.stop_fetching_inputs()
return
I built and led the Workers observability team at Cloudflare, and the habit I would carry over is simple: send these log lines to the tool you already alert from before the first long run. Modal documents integrations for Datadog and for OpenTelemetry providers. A WARN that nobody reads at hour three can easily become the restart that costs you hour twenty.
Gotchas that stop a cluster from starting or finishing
- Partial nodes are rejected.
H100:4is invalid for a cluster andH100:8is valid. CPU-only clustered functions aren’t supported. A10 nodes top out at four GPUs, but A10s don’t support RDMA anyway. - B300 needs CUDA 13.1 or later. The RDMA image example in Modal’s docs pins CUDA 12.9.1. For B300 nodes, pick a CUDA 13 base image and libraries built for it. The same applies to
gpu="B200+", which may land on a B300. - H100 requests can land on H200s. Modal may upgrade
gpu="H100"to an H200 at no extra cost. UseH100!when you benchmark and need the hardware fixed. - Ray fails on IPv6.
container_ips()returns IPv6 unless you passfamily="ipv4". - Results from other ranks disappear. Only rank 0’s return value comes back. Gather results on rank 0 first.
- Autoscaling settings that aren’t multiples of the cluster size are invalid, because those settings count nodes.
- The cluster can’t be reached from outside. i6pn addresses are private to the workspace and limited to one region. To expose a dashboard or endpoint, open a
modal.Tunnelormodal.forwardfrom one container, usually rank 0. - The decorator name differs between pages. The GA announcement and the Clusters docs use
@modal.clustered. The cluster networking page still refers to the earlier Beta name,@modal.experimental.clustered. The announcement’s snippet also callscluster.private_ips(), while the Clusters docs documentcontainer_ips(). Write new code against the Clusters docs.
A checklist before your first multi-node run
- The plan’s GPU concurrency covers nodes × 8, or you’ve talked to Modal about a larger limit.
- The GPU type supports RDMA (H100, H200, B200, B300 or GB200), and the image’s CUDA version matches it.
- The image includes
libibverbs1on a CUDA base image, and the RDMA benchmark has run once. - The leader address and node rank come from
modal.Cluster.from_context(), and Ray uses IPv4. timeout,startup_timeoutandmodal.Retriesare set, and checkpoints go to a Volume.- Rank 0 gathers results, and for inference it waits for every rank before starting the next input.
[gpu-health]log lines go to a tool someone actually watches.
Running on Modal? See how Polylane monitors Modal in production.
Common questions.
When did Modal Clusters become generally available?
Modal announced general availability on 1 October 2026, in a post by Peyton Walters on the Modal blog. Clusters are open to every workspace, and Modal says it tested them for 1.5 years before GA. You enable them with the `@modal.clustered` decorator.
Which GPUs support RDMA on Modal Clusters?
RDMA works on H100, H200, B200, B300 and GB200 GPUs. B300 clusters get 6,400 Gbps of RDMA networking per node and the others get 3,200 Gbps. Clusters without RDMA still get Modal's private i6pn network, which runs at 50 Gbps or more.
How large can a Modal Cluster be?
The docs allow up to 32 nodes and 256 GPUs per cluster, and say to contact Modal for anything larger. Your plan's GPU limits also cap cluster size. The Starter plan allows 10 concurrent GPUs, Team allows 50, and Enterprise is custom.
Can a cluster node use fewer than eight GPUs?
No. Clustered functions must request every GPU on the node, so `H100:8` is valid and `H100:4` is rejected. CPU-only clustered functions aren't supported either.
What happens if one node fails or is preempted?
If an input fails on any container, Modal stops all the others and marks the whole call as failed. A preemption also stops the entire cluster, and Modal retries it with the same input. Set `modal.Retries` and write checkpoints to a Volume, because the `nonpreemptible` option isn't supported for GPU Functions.
Do I have to configure NCCL or InfiniBand myself?
No. With `rdma=True`, Modal installs the cloud's drivers, configures the NICs and sets the NCCL environment variables. Your image still needs the userspace libraries, usually `libcudart.so`, `libibverbs.so.1` and `libmlx5.so.1`. A CUDA base image plus `apt_install("libibverbs1")` covers them.
How are Modal Clusters billed?
By the second, at the normal per-GPU rates, with no hourly minimum or reservation. The pricing page lists the B300 at $0.001972 per second and the H100 SXM5 at $0.001097 per second. Choosing a specific region multiplies base prices by 1.15 to 1.75x.
Can I serve HTTP inference from a cluster?
Yes, through a clustered Server. Web Functions aren't supported with `@modal.clustered`. A clustered Server receives traffic only on rank 0, which must pass requests to the other containers and keep them in step.
Sources
Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.
Related
- Cloudflare Durable Objects pending I/O keep-alive guide
From 2026-10-01, pending I/O keeps Cloudflare Durable Objects in memory after the client leaves. What counts, the 15-minute limit, flags and billing.
- Cloudflare K2: how to adopt serverless event streams
Cloudflare K2, announced 1 October 2026, is a serverless event stream on R2. What it is, how it differs from Queues, setup steps and beta limits.
- Turn your app into a context graph
Agents that run software need a context graph of the app: every resource, what it connects to, the repository that deploys it and the team that owns it. How we built one on Durable Objects, keep it fresh across every provider without polling AWS, decide what counts as a change, and delete from it safely.
- Amazon Bedrock Managed Agents (OpenAI preview): setup and limits
Amazon Bedrock Managed Agents, powered by OpenAI, entered preview on 29 Sep 2026. What it is, how to run a first session, and the preview limits to plan for.
- Cloudflare Containers agent sandboxes: startup and setup
Cloudflare Containers now start agent sandboxes in a median 648 ms. Set up the durable_object policy, runtime images and snapshots, and know the limits.