4.3 Single-node NCCL and NVLink Switch

NCP-AII · Cluster Test and Verification (33% of the exam) · Official objective: “Perform single-node NCCL (including verifying NVLink Switch).”

Running nccl-tests on one node, Fabric Manager and checking NVLink peer-to-peer.

Key points

  1. nccl-tests check both the performance and the correctness of NCCL operations. All-reduce is the collective most training jobs use. '-g 8' uses all 8 GPUs. '-b' and '-e' set the smallest and largest message size. '-f 2' doubles the size each step.

    What NVIDIA says (2)

    “These tests check both the performance and the correctness of NCCL operations.”

    — NVIDIA nccl-tests

    “Run on single node with 8 GPUs ( -g 8 ), scanning from 8 Bytes to 128MiB (Mebibytes), doubling between each test ( -f 2 ) : $ ./build/all_reduce_perf -b 8 -e 128M -f 2 -g 8”

    — NVIDIA nccl-tests

  2. NVLink Switch systems (HGX/DGX) need Fabric Manager (FM) to set up the NVSwitch fabric between GPUs. FM runs as the nvidia-fabricmanager service. NVIDIA says FM must be running to restore NVLink peer-to-peer after MIG mode is disabled.

    What NVIDIA says (2)

    “registers the daemon as the nvidia-fabricmanager system service.”

    — NVIDIA Fabric Manager User Guide

    “To successfully restore GPU NVLink peer-to-peer capability after the MIG mode is disabled on these systems, the FM service must be running.”

    — NVIDIA Fabric Manager User Guide

  3. Peer-to-peer (P2P) means one GPU reads or writes another GPU's memory directly. nvidia-smi topo -p2p prints a matrix of P2P status. The capability is p for PCIe and n for NVLink.

    What NVIDIA says (2)

    “You can use nvidia-smi topo -p2p <capability> to print a matrix of P2P status between GPU pairs.”

    — NCCL Troubleshooting: GPU troubleshooting

    “The <capability> value is p for PCIe and n for NVLink.”

    — NCCL Troubleshooting: GPU troubleshooting

Key terms

Try it

Sample question

You run NCCL tests on a single 8-GPU node. Which command matches NVIDIA's nccl-tests example?

Show the answer

Answer: ./build/all_reduce_perf -b 8 -e 128M -f 2 -g 8

nccl-tests check both the performance and the correctness of NCCL operations. All-reduce is the collective most training jobs use. '-g 8' uses all 8 GPUs. '-b' and '-e' set the smallest and largest message size. '-f 2' doubles the size each step.

What NVIDIA says (2)

“These tests check both the performance and the correctness of NCCL operations.”

— NVIDIA nccl-tests

“Run on single node with 8 GPUs ( -g 8 ), scanning from 8 Bytes to 128MiB (Mebibytes), doubling between each test ( -f 2 ) : $ ./build/all_reduce_perf -b 8 -e 128M -f 2 -g 8”

— NVIDIA nccl-tests

Practice 4.3 (3 questions) Full Cluster Test and Verification guide

← 4.2 Running HPL · 4.4 Signal quality on cables →