4.1 Single-node stress test

NCP-AII · Cluster Test and Verification (33% of the exam) · Official objective: “Perform a single-node stress test.”

Running the NVSM pre-flight stress test and DCGM diagnostics on one node.

Key points

  1. DCGM (Data Center GPU Manager) diagnostics come in run levels. A higher level runs longer and includes every test below it. Level 1 is the quick readiness check. Level 4 is the extended, longer-running hardware diagnostic.

    What NVIDIA says (3)

    “higher numbered tests include all beneath.”

    — DCGM User Guide: DCGM Diagnostics

    “4 - Extended (Longer-running System HW Diagnostics)”

    — DCGM User Guide: DCGM Diagnostics

    “Level 1 tests to use as a readiness metric”

    — DCGM User Guide: DCGM Diagnostics

  2. A single-node stress test checks one server alone, before testing the fabric. On DGX, NVIDIA runs it with NVSM.

    What NVIDIA says (2)

    “To run the tests, use NVSM.”

    — DGX H100/H200 Service Manual: Introduction

    “NVIDIA recommends running the pre-flight stress test before putting a system into a production environment or after servicing.”

    — DGX H100/H200 Service Manual: Introduction

Key terms

Sample question

You want DCGM's longest built-in hardware diagnostic on one node before you hand it to users. Which run level do you pick?

Show the answer

Answer: dcgmi diag -r 4

DCGM (Data Center GPU Manager) diagnostics come in run levels. A higher level runs longer and includes every test below it. Level 1 is the quick readiness check. Level 4 is the extended, longer-running hardware diagnostic.

What NVIDIA says (3)

“higher numbered tests include all beneath.”

— DCGM User Guide: DCGM Diagnostics

“4 - Extended (Longer-running System HW Diagnostics)”

— DCGM User Guide: DCGM Diagnostics

“Level 1 tests to use as a readiness metric”

— DCGM User Guide: DCGM Diagnostics

Practice 4.1 (2 questions) Full Cluster Test and Verification guide

← 3.7 NGC CLI on hosts · 4.2 Running HPL →