4.3 Single-node NCCL and NVLink Switch
Running nccl-tests on one node, Fabric Manager and checking NVLink peer-to-peer.
Key points
nccl-tests check both the performance and the correctness of NCCL operations. All-reduce is the collective most training jobs use. '-g 8' uses all 8 GPUs. '-b' and '-e' set the smallest and largest message size. '-f 2' doubles the size each step.
What NVIDIA says (2)
“These tests check both the performance and the correctness of NCCL operations.”
“Run on single node with 8 GPUs ( -g 8 ), scanning from 8 Bytes to 128MiB (Mebibytes), doubling between each test ( -f 2 ) : $ ./build/all_reduce_perf -b 8 -e 128M -f 2 -g 8”
NVLink Switch systems (HGX/DGX) need Fabric Manager (FM) to set up the NVSwitch fabric between GPUs. FM runs as the nvidia-fabricmanager service. NVIDIA says FM must be running to restore NVLink peer-to-peer after MIG mode is disabled.
What NVIDIA says (2)
“registers the daemon as the nvidia-fabricmanager system service.”
“To successfully restore GPU NVLink peer-to-peer capability after the MIG mode is disabled on these systems, the FM service must be running.”
Peer-to-peer (P2P) means one GPU reads or writes another GPU's memory directly. nvidia-smi topo -p2p prints a matrix of P2P status. The capability is p for PCIe and n for NVLink.
What NVIDIA says (2)
“You can use nvidia-smi topo -p2p <capability> to print a matrix of P2P status between GPU pairs.”
“The <capability> value is p for PCIe and n for NVLink.”
Key terms
- nccl-tests: NVIDIA's programs that measure the speed and correctness of NCCL collectives such as all_reduce.
- Fabric Manager: The service that sets up NVSwitch-based NVLink so GPUs can talk peer to peer.
- GPU peer-to-peer: GPUs reading and writing each other memory directly over NVLink or PCIe.
Try it
Sample question
You run NCCL tests on a single 8-GPU node. Which command matches NVIDIA's nccl-tests example?
Show the answer
Answer: ./build/all_reduce_perf -b 8 -e 128M -f 2 -g 8
nccl-tests check both the performance and the correctness of NCCL operations. All-reduce is the collective most training jobs use. '-g 8' uses all 8 GPUs. '-b' and '-e' set the smallest and largest message size. '-f 2' doubles the size each step.
What NVIDIA says (2)
“These tests check both the performance and the correctness of NCCL operations.”
“Run on single node with 8 GPUs ( -g 8 ), scanning from 8 Bytes to 128MiB (Mebibytes), doubling between each test ( -f 2 ) : $ ./build/all_reduce_perf -b 8 -e 128M -f 2 -g 8”
Practice 4.3 (3 questions) Full Cluster Test and Verification guide