4.10 NCCL east-west bandwidth

NCP-AII · Cluster Test and Verification (33% of the exam) · Official objective: “Run NCCL to verify E/W fabric bandwidth.”

Multi-node nccl-tests with MPI and reading busbw.

Key points

  1. East-west (E/W) traffic flows between compute nodes over the fabric. This run starts 64 MPI processes, one per GPU, across 8 nodes. MPI (Message Passing Interface) is the standard for launching multi-node jobs. The binaries must be built with MPI=1.

    What NVIDIA says (3)

    “Run 64 MPI processes on nodes with 8 GPUs each, for a total of 64 GPUs spread across 8 nodes.”

    — NVIDIA nccl-tests

    “mpirun -np 64 -N 8 ./build/all_reduce_perf -b 8 -e 8G -f 2 -g 1”

    — NVIDIA nccl-tests

    “(NB: The nccl-tests binaries must be compiled with MPI=1 for this case)”

    — NVIDIA nccl-tests

  2. algbw (algorithm bandwidth) is data size divided by time. busbw (bus bandwidth) corrects for the collective's traffic pattern, so it can be compared with hardware link speed. The nccl-tests README points to the busbw column.

    What NVIDIA says (1)

    “See the Performance page for explanation about numbers, and in particular the "busbw" column.”

    — NVIDIA nccl-tests

  3. Multi-node runs start one process per GPU with MPI. The test binaries must be built with MPI support.

    What NVIDIA says (1)

    “(NB: The nccl-tests binaries must be compiled with MPI=1 for this case)”

    — NVIDIA nccl-tests

Key terms

Try it

Sample question

You want nccl-tests to measure east-west fabric bandwidth across 8 nodes with 8 GPUs each. Which command matches NVIDIA's example?

Show the answer

Answer: mpirun -np 64 -N 8 ./build/all_reduce_perf -b 8 -e 8G -f 2 -g 1

East-west (E/W) traffic flows between compute nodes over the fabric. This run starts 64 MPI processes, one per GPU, across 8 nodes. MPI (Message Passing Interface) is the standard for launching multi-node jobs. The binaries must be built with MPI=1.

What NVIDIA says (3)

“Run 64 MPI processes on nodes with 8 GPUs each, for a total of 64 GPUs spread across 8 nodes.”

— NVIDIA nccl-tests

“mpirun -np 64 -N 8 ./build/all_reduce_perf -b 8 -e 8G -f 2 -g 1”

— NVIDIA nccl-tests

“(NB: The nccl-tests binaries must be compiled with MPI=1 for this case)”

— NVIDIA nccl-tests

Practice 4.10 (3 questions) Full Cluster Test and Verification guide

← 4.9 ClusterKit node assessment · 4.11 NCCL burn-in →