4.12 HPL burn-in

NCP-AII · Cluster Test and Verification (33% of the exam) · Official objective: “Perform HPL burn-in.”

Using HPL as a math-heavy load for burn-in with the sample Slurm scripts.

Key points

  1. HPL (High-Performance Linpack) solves a large dense linear system. It loads the GPUs hard and also uses the network. NVIDIA's SuperPOD admin guide lists NCCL for the fabric and HPL for math-intensive applications with network communication.

    What NVIDIA says (2)

    “Math intensive applications with network communications”

    — DGX SuperPOD Administration Guide: System Health Checks and Debugging

    “The NVIDIA HPL benchmark expects one GPU per MPI process.”

    — NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

  2. HPL.dat is HPL's input file. It sets the problem size and process grid. NVIDIA ships samples of both input files and Slurm batch-job scripts with the package.

    What NVIDIA says (2)

    “Samples of Slurm batch-job scripts in sample-slurm directory”

    — NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

    “Samples of input files in sample-dat directory”

    — NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

Key terms

Sample question

Why is HPL used for burn-in alongside NCCL?

Show the answer

Answer: HPL stresses math-heavy compute with network communication. NCCL targets the fabric.

HPL (High-Performance Linpack) solves a large dense linear system. It loads the GPUs hard and also uses the network. NVIDIA's SuperPOD admin guide lists NCCL for the fabric and HPL for math-intensive applications with network communication.

What NVIDIA says (2)

“Math intensive applications with network communications”

— DGX SuperPOD Administration Guide: System Health Checks and Debugging

“The NVIDIA HPL benchmark expects one GPU per MPI process.”

— NVIDIA HPC Benchmarks: NVIDIA HPL Benchmark

Practice 4.12 (2 questions) Full Cluster Test and Verification guide

← 4.11 NCCL burn-in · 4.13 NeMo burn-in →