NCCL on 2–4 GB10 Machines: Testing the Link
NCCL is NVIDIA's library through which the GPUs of different machines exchange data during distributed work. Here we build it on every machine and run a bandwidth test of the link — before you start a real workload. We follow the NVIDIA playbook "NCCL for Multiple Sparks".
all_gather_perf through mpirun with the network variables from the playbook) and the quick path with NVIDIA's scripts. We removed: all speed and latency figures and the "expected output" (not proven), the gradient arithmetic, the "2.5 times faster" promises, the sample in-house model and internal examples of machine counts, the fixed addresses, the "gradient compression" task and the large training example. Only a small AllReduce example remains — not run.
01What you'll learn
- What NCCL is and what its main operations are.
- How the machines are connected and why the network is done before this lesson.
- How to build NCCL and its test on every machine — the quick way and by hand.
- How to run the test through
mpirunand how to read the result. - What to do if the test hangs or cannot find an interface.
- What the smallest AllReduce example looks like.
02Before you start
- Two, three or four DGX Spark systems (or of the same class — see the note above).
- Connected following the matching NVIDIA playbook: two machines — Connect Two Sparks; three — the lesson "Three GB10 Machines in a Ring"; four — Connect Multiple DGX Spark through a Switch. It covers the cables, the addresses and passwordless SSH.
- A driver installed (
nvidia-smi), the CUDA toolkit available (nvcc --version) andsudorights (sudo whoami). - Internet access to download code from GitHub.
| What | Value (per the NVIDIA playbook) |
|---|---|
| Machines | 2, 3 or 4 |
| Time | about 30 minutes for setup and validation |
| Risk | medium — involves network changes |
| Rollback | deleting the NCCL and test folders |
| NCCL | version v2.30.7-1, built from source |
03Steps
-
What NCCL is
NCCL (pronounced "nickel") is a library for collective communication between GPUs on many machines. "Collective" means all of them take part at once, rather than one sending to another. The main operations:
Operation What it does AllReduce combines (for example sums) values from all and returns the result to all AllGather each sends its piece and each receives all the pieces Broadcast one sends to all ReduceScatter combines and distributes the parts among the participants Send/Recv direct transfer between two GPUs The test in this lesson uses AllGather, because that is what the NVIDIA playbook does.
-
The network — first
Do the connection for your number of machines following its playbook (see "Before you start"): cables, addresses, passwordless SSH and a connectivity check. According to a note in the playbook, full bandwidth can be achieved with just one QSFP cable; when two cables are connected, all four interfaces must be assigned an address to get full bandwidth.
-
Quick path: NVIDIA's scripts
The playbook offers two helper scripts that do the build and the test for you. They run from Machine 1 and take the addresses you connect to over SSH (the so-called management addresses). Read the script before you run it — it executes commands on all your machines.
bash · from Machine 1curl -fsSL https://raw.githubusercontent.com/NVIDIA/dgx-spark-playbooks/refs/heads/main/nvidia/nccl/assets/setup.sh -o setup.sh curl -fsSL https://raw.githubusercontent.com/NVIDIA/dgx-spark-playbooks/refs/heads/main/nvidia/nccl/assets/launch.sh -o launch.sh # builds NCCL v2.30.7-1 and the test on all machines (two machines): bash setup.sh <address-of-machine-2> # runs the test (two machines, direct cable): bash launch.sh --topology direct <address-of-machine-1> <address-of-machine-2>For three machines:
bash setup.sh <machine-2> <machine-3>and--topology ring. For four: one more address and--topology switch. The script expects the wired management interfaceenP7s7; on Wi-Fi instead of a cable, prefix the command withMGMT_IFNAME=<wifi-interface>(see step 6). If you want to understand it or hunt for an error — follow the manual steps below. -
By hand: build NCCL (on every machine)
Why from source? The playbook builds NCCL with support for the Blackwell architecture (
sm_121), which is the one in GB10.bash · on every machinesudo apt-get update && sudo apt-get install -y libopenmpi-dev git clone -b v2.30.7-1 https://github.com/NVIDIA/nccl.git ~/nccl/ cd ~/nccl/ make -j src.build NVCC_GENCODE="-gencode=arch=compute_121,code=sm_121" export CUDA_HOME="/usr/local/cuda" export MPI_HOME="/usr/lib/aarch64-linux-gnu/openmpi" export NCCL_HOME="$HOME/nccl/build/" export LD_LIBRARY_PATH="$NCCL_HOME/lib:$CUDA_HOME/lib64/:$MPI_HOME/lib:$LD_LIBRARY_PATH" -
By hand: build the test (on every machine)
bash · on every machinegit clone https://github.com/NVIDIA/nccl-tests.git ~/nccl-tests/ cd ~/nccl-tests/ make MPI=1 -
Check the ports and find the addresses
bash · on every machineibdev2netdev ip addr show enP7s7The first command shows the ConnectX-7 ports and whether they are "(Up)" — see the connection lesson for the expected count. The second shows the machine's management address — the one you connect to over SSH. Write it down for every machine. If your machines are on Wi-Fi rather than a cable, replace
enP7s7with the Wi-Fi interface (the name is shown byip -o link show) and use the Wi-Fi addresses. On all machines the interface must be the same (per the playbook). -
Run the test
From Machine 1 (it is the "launcher" —
mpirunstarts the test on the others over SSH). Replace the addresses with yours. Each address is followed by:1— one task per machine. An example for two machines:bash · from Machine 1export UCX_NET_DEVICES=enP7s7 export NCCL_SOCKET_IFNAME=enP7s7 export OMPI_MCA_btl_tcp_if_include=enP7s7 mpirun -np 2 -H <address-of-machine-1>:1,<address-of-machine-2>:1 \ --mca plm_rsh_agent "ssh -o UserKnownHostsFile=/dev/null -o StrictHostKeyChecking=no" \ -x LD_LIBRARY_PATH=$LD_LIBRARY_PATH \ $HOME/nccl-tests/build/all_gather_perfFor a larger buffer that loads more of the 200 Gbps line, add
-b 16G -e 16G -f 2at the end. For three machines:-np 3, the three addresses and — for a ring only — two more variables beforempirun:export NCCL_IB_SUBNET_AWARE_ROUTING=1andexport NCCL_NET_PLUGIN=none. For four machines:-np 4and the four addresses, with no extra variables.⚠️What the SSH option doesTheStrictHostKeyChecking=noline is from the playbook and saves the question on the first connection. It weakens the check of the machines — use it only in a closed lab network that you trust. -
How to read the result
The test prints a table for different data sizes. In the standard
nccl-testsoutput, the columns that matter are time,algbwandbusbw(data rate) and#wrong, which must be 0. We give no concrete numbers here — we have not run it. To compare with what to expect, see NVIDIA's performance guide. -
If something does not work
Problem Cause What to do (per NVIDIA) mpirunhangs or times outan SSH problem Check ssh <address>— it must work without a password. Trympirun -np 2 -H <address-1>:1,<address-2>:1 hostname. Check the keys on all machines.Interface not found wrong name or it is down See ibdev2netdevand whether the interface has an address.NCCL build fails OpenMPI is missing or CUDA is not the same Check CUDA and the required libraries. -
Clean up
bash · on every machinerm -rf ~/nccl/ rm -rf ~/nccl-tests/This removes only the code from this lesson. The network settings are reverted following the connection playbook.
-
The smallest AllReduce example (optional)
Once the link is checked, see the same operation through PyTorch. Each machine makes a number, and AllReduce sums them and returns the total to everyone. With two machines we expect
[3.0, 3.0, 3.0]on each (1 + 2). ⚠️ Not run; you need a PyTorch with support for the GB10 GPU — see the lesson "PyTorch on GB10".python · allreduce_demo.pyimport torch import torch.distributed as dist dist.init_process_group("nccl") rank = dist.get_rank() torch.cuda.set_device(0) # one GPU per machine x = torch.tensor([float(rank + 1)] * 3, device="cuda") dist.all_reduce(x, op=dist.ReduceOp.SUM) print(f"machine {rank}: {x.tolist()}") dist.destroy_process_group()bash · on every machine (with its own node_rank)NCCL_SOCKET_IFNAME=enP7s7 torchrun \ --nnodes=2 --nproc_per_node=1 \ --node_rank=<0 or 1> \ --master_addr=<address-of-machine-1> --master_port=29500 \ allreduce_demo.py
04Check
- The connection for your number of machines is done following its playbook and SSH between them works without a password.
nvidia-smi,nvcc --versionandsudowork on every machine.- NCCL and the test are built on every machine (or
setup.shhas completed). ibdev2netdevshows the expected ports as "(Up)".all_gather_perffinishes without errors across all machines.- You know how to clean up.
Quiz
1. What is NCCL?
2. What connects the machines to each other in NVIDIA's playbook?
3. What must be ready before you build NCCL?
4. What do you check if mpirun hangs?
05What's next
06Sources
Pages were opened on 03.10.2026.
- NVIDIA: NCCL for Multiple Sparks 🌐 global — overview, requirements, time and risk (last updated per NVIDIA: 15.12.2025).
- two machines · three machines · four machines · troubleshooting — the steps and the table of issues.
- GitHub: setup.sh and launch.sh scripts — the quick path and the description of the topologies.
- GitHub: NCCL · GitHub: nccl-tests — code and releases.
- NVIDIA: NCCL user guide — the operations and settings.