4.11 NCCL burn-in
Running nccl-tests in long loops with correctness checks, and triaging differences.
Key points
A burn-in runs a load for a long time to catch parts that fail under sustained stress. -N sets run cycles, and 0 means run forever. -c checks results, which is slower on many GPUs.
What NVIDIA says (2)
“-N,--run_cycles <cycle count> run & print each cycle. Default : 1; 0=infinite.”
“-c,--check <check iteration count> perform count iterations, checking correctness of results on each iteration.”
NCCL is NVIDIA's GPU communication library. NVIDIA names it the tool for validating the fabric. A repeatable gap on the same hardware points to a component. Suspect nodes leave the batch partition for triage.
What NVIDIA says (2)
“However, over multiple runs on the same sets of hardware a difference is found, it can indicate an issue with some component of that system.”
“If an issue is found, it should be removed from the batch partition for initial triage.”
Key terms
- nccl-tests: NVIDIA's programs that measure the speed and correctness of NCCL collectives such as all_reduce.
- Burn-in: Running a heavy workload for a long time to expose weak parts before production.
Sample question
For an NCCL burn-in you want all_reduce_perf to keep cycling and check results for correctness. Which options do this?
Show the answer
Answer: -N 0 for infinite run cycles, and -c to check correctness each iteration.
A burn-in runs a load for a long time to catch parts that fail under sustained stress. -N sets run cycles, and 0 means run forever. -c checks results, which is slower on many GPUs.
What NVIDIA says (2)
“-N,--run_cycles <cycle count> run & print each cycle. Default : 1; 0=infinite.”
“-c,--check <check iteration count> perform count iterations, checking correctness of results on each iteration.”
Practice 4.11 (2 questions) Full Cluster Test and Verification guide