
We’re introducing InfiniBand clusters in SF Autoresearch. Create a group of dedicated H100 machines, connect them over a private InfiniBand network, and run distributed training from your terminal or agent.
Clusters extend Autoresearch to experiments that need GPUs on several machines working together. Each node contributes eight H100 GPUs. InfiniBand RDMA provides the connection between machines for NCCL operations such as exchanging gradients during training.
Each cluster gets a fresh Elastic InfiniBand Partition. Autoresearch places the nodes on the same fabric and manages their partition membership. You use the same commands to run code, copy files, and read logs that you use on an individual node.
Clusters are currently in internal preview. Contact us to register your interest in customer access.
On an account with cluster access enabled, create two nodes with 16 H100 GPUs:
autor cluster create --name train \
--chip h100-8 --nodes 2 --gpus 16 \
--image pytorch-2.13-cuda12.9 --waitThis creates train-0 and train-1 in their own InfiniBand partition. The --wait flag waits until both nodes run. The image includes PyTorch and CUDA for the training example below.
Hold both nodes during setup, then inspect the cluster:
for rank in 0 1; do
autor hold "train-$rank" --until 1h
done
autor cluster show trainA hold keeps an idle node awake and bills at its normal rate.
Install the InfiniBand libraries on both nodes so NCCL can use RDMA. Then copy your existing PyTorch distributed training script, train.py, to each node:
for rank in 0 1; do
autor run "train-$rank" --timeout 300 -- sh -c \
'sudo apt-get update && sudo apt-get install -y ibverbs-utils ibverbs-providers'
autor cp ./train.py "train-$rank:~/train.py"
doneAutoresearch supplies each node’s rank, cluster size, and rendezvous address. Start eight training processes per node:
for rank in 0 1; do
autor run "train-$rank" -d -- sh -c '
NCCL_DEBUG=INFO torchrun \
--nnodes "$GMN_CLUSTER_SIZE" \
--node-rank "$GMN_CLUSTER_RANK" \
--nproc-per-node 8 \
--master-addr "$GMN_CLUSTER_MASTER_ADDR" \
--master-port 29500 \
"$HOME/train.py"
'
done
autor logs train-1
autor logs train-0 --followThe detached launches let both nodes start without waiting for the other command to finish. Autoresearch supplies the private network address for rendezvous, and NCCL moves training data over InfiniBand. With NCCL_DEBUG=INFO, check the logs for data channels using NET/IB.
In our October 2–3 production validation, two eight-GPU H100 nodes joined a fresh partition and ran NCCL all-reduce across all 16 GPUs. The run reached 112.6 GB/s of NCCL bus bandwidth with a 1 GiB payload.
| Measurement | Result |
|---|---|
| NCCL all-reduce bus bandwidth, 16 GPUs, 1 GiB | 112.6 GB/s |
| NCCL all-reduce algorithm bandwidth | 60.1 GB/s |
| RDMA write bandwidth, one port, 1 GiB memory region | 370.4 Gb/s (46.3 GB/s) |
All eight InfiniBand adapters on each machine carried traffic. NCCL’s data channels used NET/IB, with no NET/Socket fallback.
Algorithm bandwidth measures payload size divided by operation time. Bus bandwidth adjusts that figure for the communication required by all-reduce. NVIDIA’s guide explains the distinction.
Each member bills at the ordinary eight-GPU node rate from the moment it is ready, with no minimum charge. You do not pay while the cluster waits for capacity or prepares its machines. Deleting the cluster stops billing without waiting for the machines to leave the partition.
Our 38-minute validation run cost about $27 at that test’s Friday-discounted rate of $0.71 per minute for both nodes.
Copy your results out before deleting the cluster. Its nodes and their disks are deleted together:
autor cluster rm trainA cluster runs as a group. If one node stops or is lost, the whole cluster ends and its disks are deleted. Save checkpoints outside the cluster for runs you need to resume.
Read the cluster documentation for the full workflow, or contact us about customer access.
