4.12 HPL burn-in
Using HPL as a math-heavy load for burn-in with the sample Slurm scripts.
Key points
HPL (High-Performance Linpack) solves a large dense linear system. It loads the GPUs hard and also uses the network. NVIDIA's SuperPOD admin guide lists NCCL for the fabric and HPL for math-intensive applications with network communication.
What NVIDIA says (2)
“Math intensive applications with network communications”
“The NVIDIA HPL benchmark expects one GPU per MPI process.”
HPL.dat is HPL's input file. It sets the problem size and process grid. NVIDIA ships samples of both input files and Slurm batch-job scripts with the package.
What NVIDIA says (2)
“Samples of Slurm batch-job scripts in sample-slurm directory”
“Samples of input files in sample-dat directory”
Key terms
- High-Performance Linpack: A math-heavy benchmark used to load and compare systems; NVIDIA HPL runs one GPU per MPI process.
- Burn-in: Running a heavy workload for a long time to expose weak parts before production.
Sample question
Why is HPL used for burn-in alongside NCCL?
Show the answer
Answer: HPL stresses math-heavy compute with network communication. NCCL targets the fabric.
HPL (High-Performance Linpack) solves a large dense linear system. It loads the GPUs hard and also uses the network. NVIDIA's SuperPOD admin guide lists NCCL for the fabric and HPL for math-intensive applications with network communication.
What NVIDIA says (2)
“Math intensive applications with network communications”
“The NVIDIA HPL benchmark expects one GPU per MPI process.”
Practice 4.12 (2 questions) Full Cluster Test and Verification guide