4.11 NCCL burn-in

NCP-AII · Cluster Test and Verification (33% of the exam) · Official objective: “Perform NCCL burn-in.”

Running nccl-tests in long loops with correctness checks, and triaging differences.

Key points

  1. A burn-in runs a load for a long time to catch parts that fail under sustained stress. -N sets run cycles, and 0 means run forever. -c checks results, which is slower on many GPUs.

    What NVIDIA says (2)

    “-N,--run_cycles <cycle count> run & print each cycle. Default : 1; 0=infinite.”

    — NVIDIA nccl-tests

    “-c,--check <check iteration count> perform count iterations, checking correctness of results on each iteration.”

    — NVIDIA nccl-tests

  2. NCCL is NVIDIA's GPU communication library. NVIDIA names it the tool for validating the fabric. A repeatable gap on the same hardware points to a component. Suspect nodes leave the batch partition for triage.

    What NVIDIA says (2)

    “However, over multiple runs on the same sets of hardware a difference is found, it can indicate an issue with some component of that system.”

    — DGX SuperPOD Administration Guide: System Health Checks and Debugging

    “If an issue is found, it should be removed from the batch partition for initial triage.”

    — DGX SuperPOD Administration Guide: System Health Checks and Debugging

Key terms

Sample question

For an NCCL burn-in you want all_reduce_perf to keep cycling and check results for correctness. Which options do this?

Show the answer

Answer: -N 0 for infinite run cycles, and -c to check correctness each iteration.

A burn-in runs a load for a long time to catch parts that fail under sustained stress. -N sets run cycles, and 0 means run forever. -c checks results, which is slower on many GPUs.

What NVIDIA says (2)

“-N,--run_cycles <cycle count> run & print each cycle. Default : 1; 0=infinite.”

— NVIDIA nccl-tests

“-c,--check <check iteration count> perform count iterations, checking correctness of results on each iteration.”

— NVIDIA nccl-tests

Practice 4.11 (2 questions) Full Cluster Test and Verification guide

← 4.10 NCCL east-west bandwidth · 4.12 HPL burn-in →