4.2 Running HPL
How NVIDIA HPL maps MPI processes to GPUs and where its sample inputs live.
Key points
HPL (High-Performance Linpack) solves a large dense linear system to stress GPUs and the network. MPI (Message Passing Interface) starts one process per rank. NVIDIA's HPL expects one GPU per MPI process, so the process count equals the GPU count.
What NVIDIA says (1)
“The NVIDIA HPL benchmark expects one GPU per MPI process. As such, set the number of MPI processes to match the number of available GPUs in the cluster.”
An MPI process, or rank, is one copy of the program. NVIDIA HPL maps one GPU to each rank.
What NVIDIA says (1)
“The NVIDIA HPL benchmark expects one GPU per MPI process.”
Key terms
- High-Performance Linpack: A math-heavy benchmark used to load and compare systems; NVIDIA HPL runs one GPU per MPI process.
Sample question
How many MPI processes should you start for the NVIDIA HPL benchmark on 8 nodes with 8 GPUs each?
Show the answer
Answer: 64, one per GPU
HPL (High-Performance Linpack) solves a large dense linear system to stress GPUs and the network. MPI (Message Passing Interface) starts one process per rank. NVIDIA's HPL expects one GPU per MPI process, so the process count equals the GPU count.
What NVIDIA says (1)
“The NVIDIA HPL benchmark expects one GPU per MPI process. As such, set the number of MPI processes to match the number of available GPUs in the cluster.”
Practice 4.2 (2 questions) Full Cluster Test and Verification guide
← 4.1 Single-node stress test · 4.3 Single-node NCCL and NVLink Switch →