The KAGAMI mark КАГАМИ
kagami.bg/academy · lesson · machine-readable viewUPDATED 2026-10-03
IDENTITY
module
GX10-04-46 · NCCL on two to four GB10-class machines: build and test the link
series
GX10 (local AI server class: NVIDIA GB10, e.g. ASUS Ascent GX10 / DGX Spark)
level
Advanced
duration
about 30 min for setup and validation (per NVIDIA); risk level medium
prerequisites
Two, three or four DGX Spark systems already connected following the matching NVIDIA connection playbook (two: direct cable; three: ring; four: switch), working passwordless SSH, nvidia-smi, nvcc, sudo
trust_label
UPDATED 2026-10-03 (checked against the NVIDIA playbook "NCCL for Multiple Sparks", last updated 2025-12-15 per NVIDIA, and its helper scripts on GitHub). NOT TESTED: no cluster hardware was available; no command was run. The playbook targets DGX Spark; ASUS Ascent GX10 is the same GB10 class but was not checked.
versions_seen
NCCL v2.30.7-1 (pinned by the playbook) · NCCL latest tag on GitHub on 2026-10-03: v2.32.3-1 (not used by the playbook, not checked on GB10)
language
human view: english edition · bulgarian edition: /academy/gx10/ (same file name)
previous / next
GX10 series index / GX10 series index
PURPOSE

Build NCCL from source with Blackwell support on every node, build nccl-tests, and run the all_gather_perf bandwidth test across the nodes with mpirun, following the NVIDIA playbook. Learn what NCCL collectives are, what a table of results contains, and how to troubleshoot the common failures. Includes a tiny AllReduce demo for torchrun (not run).

KEY CONCEPTS
COMMANDS / PATHS
CHECKLIST
NEXT MODULE

Series index: kagami.bg/en/academy/gx10/ · related lessons: 04-204 (three machines in a ring), 04-21 (PyTorch on GB10) · NVIDIA playbooks: Connect Two Sparks, Connect Three DGX Spark in a Ring Topology, Connect Multiple DGX Spark through a Switch · offer: Quick experiment (kagami.bg/en/stalbata/)

SOURCES
TAGS
gx10nvidia-gb10ncclmpiclusterdistributed-trainingconnectx-7
UPDATED · 03.10.2026

NCCL on 2–4 GB10 Machines: Testing the Link

NCCL is NVIDIA's library through which the GPUs of different machines exchange data during distributed work. Here we build it on every machine and run a bandwidth test of the link — before you start a real workload. We follow the NVIDIA playbook "NCCL for Multiple Sparks".

⏱ ~30 min Advanced GX10 2–4 machines · QSFP · 200 Gbps NCCL · MPI · nccl-tests
NCCL, nccl-tests, MPI🔒 local Code and scripts from GitHub🌐 global
🔄
UPDATED · 03.10.2026 — what changed
The lesson was rebuilt entirely against NVIDIA's current playbook (last updated per NVIDIA: 15.12.2025). We corrected: the link between the machines (it runs over the ConnectX-7 ports and QSFP cables, not an "NVLink between units"), the way of building (NCCL is built from source with Blackwell support, not installed as a ready package), the test (all_gather_perf through mpirun with the network variables from the playbook) and the quick path with NVIDIA's scripts. We removed: all speed and latency figures and the "expected output" (not proven), the gradient arithmetic, the "2.5 times faster" promises, the sample in-house model and internal examples of machine counts, the fixed addresses, the "gradient compression" task and the large training example. Only a small AllReduce example remains — not run.
⚠️
What we have not run ourselves
We had no machines and cables, so nothing in this lesson has been executed, and there is no "TESTED" label. The playbook is for DGX Spark; ASUS Ascent GX10 is the same GB10 class, but we have not checked it against it. We have not measured speed. The playbook pins NCCL to version v2.30.7-1; a newer version (v2.32.3-1 on 03.10.2026) exists, but we have not checked it on GB10.

01What you'll learn

02Before you start

WhatValue (per the NVIDIA playbook)
Machines2, 3 or 4
Timeabout 30 minutes for setup and validation
Riskmedium — involves network changes
Rollbackdeleting the NCCL and test folders
NCCLversion v2.30.7-1, built from source
💡
Easy to get wrong
Between the machines, data travels over the ConnectX-7 ports through QSFP cables (200 Gbps per the playbook). An older version of this lesson claimed there is "NVLink" between machines with fixed speeds — that is not in NVIDIA's playbook, and we removed it.

03Steps

  1. What NCCL is

    NCCL (pronounced "nickel") is a library for collective communication between GPUs on many machines. "Collective" means all of them take part at once, rather than one sending to another. The main operations:

    OperationWhat it does
    AllReducecombines (for example sums) values from all and returns the result to all
    AllGathereach sends its piece and each receives all the pieces
    Broadcastone sends to all
    ReduceScattercombines and distributes the parts among the participants
    Send/Recvdirect transfer between two GPUs

    The test in this lesson uses AllGather, because that is what the NVIDIA playbook does.

  2. The network — first

    Do the connection for your number of machines following its playbook (see "Before you start"): cables, addresses, passwordless SSH and a connectivity check. According to a note in the playbook, full bandwidth can be achieved with just one QSFP cable; when two cables are connected, all four interfaces must be assigned an address to get full bandwidth.

  3. Quick path: NVIDIA's scripts

    The playbook offers two helper scripts that do the build and the test for you. They run from Machine 1 and take the addresses you connect to over SSH (the so-called management addresses). Read the script before you run it — it executes commands on all your machines.

    bash · from Machine 1
    curl -fsSL https://raw.githubusercontent.com/NVIDIA/dgx-spark-playbooks/refs/heads/main/nvidia/nccl/assets/setup.sh -o setup.sh
    curl -fsSL https://raw.githubusercontent.com/NVIDIA/dgx-spark-playbooks/refs/heads/main/nvidia/nccl/assets/launch.sh -o launch.sh
    
    # builds NCCL v2.30.7-1 and the test on all machines (two machines):
    bash setup.sh <address-of-machine-2>
    
    # runs the test (two machines, direct cable):
    bash launch.sh --topology direct <address-of-machine-1> <address-of-machine-2>

    For three machines: bash setup.sh <machine-2> <machine-3> and --topology ring. For four: one more address and --topology switch. The script expects the wired management interface enP7s7; on Wi-Fi instead of a cable, prefix the command with MGMT_IFNAME=<wifi-interface> (see step 6). If you want to understand it or hunt for an error — follow the manual steps below.

  4. By hand: build NCCL (on every machine)

    Why from source? The playbook builds NCCL with support for the Blackwell architecture (sm_121), which is the one in GB10.

    bash · on every machine
    sudo apt-get update && sudo apt-get install -y libopenmpi-dev
    git clone -b v2.30.7-1 https://github.com/NVIDIA/nccl.git ~/nccl/
    cd ~/nccl/
    make -j src.build NVCC_GENCODE="-gencode=arch=compute_121,code=sm_121"
    
    export CUDA_HOME="/usr/local/cuda"
    export MPI_HOME="/usr/lib/aarch64-linux-gnu/openmpi"
    export NCCL_HOME="$HOME/nccl/build/"
    export LD_LIBRARY_PATH="$NCCL_HOME/lib:$CUDA_HOME/lib64/:$MPI_HOME/lib:$LD_LIBRARY_PATH"
  5. By hand: build the test (on every machine)

    bash · on every machine
    git clone https://github.com/NVIDIA/nccl-tests.git ~/nccl-tests/
    cd ~/nccl-tests/
    make MPI=1
  6. Check the ports and find the addresses

    bash · on every machine
    ibdev2netdev
    ip addr show enP7s7

    The first command shows the ConnectX-7 ports and whether they are "(Up)" — see the connection lesson for the expected count. The second shows the machine's management address — the one you connect to over SSH. Write it down for every machine. If your machines are on Wi-Fi rather than a cable, replace enP7s7 with the Wi-Fi interface (the name is shown by ip -o link show) and use the Wi-Fi addresses. On all machines the interface must be the same (per the playbook).

  7. Run the test

    From Machine 1 (it is the "launcher" — mpirun starts the test on the others over SSH). Replace the addresses with yours. Each address is followed by :1 — one task per machine. An example for two machines:

    bash · from Machine 1
    export UCX_NET_DEVICES=enP7s7
    export NCCL_SOCKET_IFNAME=enP7s7
    export OMPI_MCA_btl_tcp_if_include=enP7s7
    
    mpirun -np 2 -H <address-of-machine-1>:1,<address-of-machine-2>:1 \
      --mca plm_rsh_agent "ssh -o UserKnownHostsFile=/dev/null -o StrictHostKeyChecking=no" \
      -x LD_LIBRARY_PATH=$LD_LIBRARY_PATH \
      $HOME/nccl-tests/build/all_gather_perf

    For a larger buffer that loads more of the 200 Gbps line, add -b 16G -e 16G -f 2 at the end. For three machines: -np 3, the three addresses and — for a ring only — two more variables before mpirun: export NCCL_IB_SUBNET_AWARE_ROUTING=1 and export NCCL_NET_PLUGIN=none. For four machines: -np 4 and the four addresses, with no extra variables.

    ⚠️
    What the SSH option does
    The StrictHostKeyChecking=no line is from the playbook and saves the question on the first connection. It weakens the check of the machines — use it only in a closed lab network that you trust.
  8. How to read the result

    The test prints a table for different data sizes. In the standard nccl-tests output, the columns that matter are time, algbw and busbw (data rate) and #wrong, which must be 0. We give no concrete numbers here — we have not run it. To compare with what to expect, see NVIDIA's performance guide.

  9. If something does not work

    ProblemCauseWhat to do (per NVIDIA)
    mpirun hangs or times outan SSH problemCheck ssh <address> — it must work without a password. Try mpirun -np 2 -H <address-1>:1,<address-2>:1 hostname. Check the keys on all machines.
    Interface not foundwrong name or it is downSee ibdev2netdev and whether the interface has an address.
    NCCL build failsOpenMPI is missing or CUDA is not the sameCheck CUDA and the required libraries.
  10. Clean up

    bash · on every machine
    rm -rf ~/nccl/
    rm -rf ~/nccl-tests/

    This removes only the code from this lesson. The network settings are reverted following the connection playbook.

  11. The smallest AllReduce example (optional)

    Once the link is checked, see the same operation through PyTorch. Each machine makes a number, and AllReduce sums them and returns the total to everyone. With two machines we expect [3.0, 3.0, 3.0] on each (1 + 2). ⚠️ Not run; you need a PyTorch with support for the GB10 GPU — see the lesson "PyTorch on GB10".

    python · allreduce_demo.py
    import torch
    import torch.distributed as dist
    
    dist.init_process_group("nccl")
    rank = dist.get_rank()
    torch.cuda.set_device(0)          # one GPU per machine
    
    x = torch.tensor([float(rank + 1)] * 3, device="cuda")
    dist.all_reduce(x, op=dist.ReduceOp.SUM)
    print(f"machine {rank}: {x.tolist()}")
    
    dist.destroy_process_group()
    bash · on every machine (with its own node_rank)
    NCCL_SOCKET_IFNAME=enP7s7 torchrun \
      --nnodes=2 --nproc_per_node=1 \
      --node_rank=<0 or 1> \
      --master_addr=<address-of-machine-1> --master_port=29500 \
      allreduce_demo.py

04Check

Quiz

1. What is NCCL?

2. What connects the machines to each other in NVIDIA's playbook?

3. What must be ready before you build NCCL?

4. What do you check if mpirun hangs?

05What's next

06Sources

Pages were opened on 03.10.2026.

  1. NVIDIA: NCCL for Multiple Sparks 🌐 global — overview, requirements, time and risk (last updated per NVIDIA: 15.12.2025).
  2. two machines · three machines · four machines · troubleshooting — the steps and the table of issues.
  3. GitHub: setup.sh and launch.sh scripts — the quick path and the description of the topologies.
  4. GitHub: NCCL · GitHub: nccl-tests — code and releases.
  5. NVIDIA: NCCL user guide — the operations and settings.