3.5 System management tools

NCP-AIO · Workload Management (23% of the exam) · Official objective: “Use system management tools to troubleshoot issues”

DCGM and nvidia-smi for finding and clearing GPU problems.

Key points

  1. DCGM (Data Center GPU Manager) diagnostics run in levels. Higher levels take longer and test more.

    What NVIDIA says (1)

    “Level 1 tests to use as a readiness metric Level 2 tests to use as an epilogue on failure Level 3 and Level 4 tests to be run by an administrator as post-mortem”

    — DCGM Diagnostics

  2. Discovery is the first check. If a GPU is missing here, deeper tests will not help.

    What NVIDIA says (1)

    “You should see a listing of all supported GPUs (and any NVSwitches) found in the system: $ dcgmi discovery -l”

    — DCGM User Guide: Getting Started

  3. dmon is device monitoring. It prints a compact line per cycle.

    What NVIDIA says (1)

    “This tool allows the user to see one line of monitoring data per monitoring cycle.”

    — nvidia-smi documentation

  4. A GPU reset clears hardware and software state on the GPU. Use -i to target one GPU.

    What NVIDIA says (1)

    “Can be used to clear GPU HW and SW state in situations that would otherwise require a machine reboot. Typically useful if a double bit ECC error has occurred.”

    — nvidia-smi documentation

Key terms

Try it

Sample question

Which DCGM diagnostic level is meant as a quick readiness check?

Show the answer

Answer: Level 1 (dcgmi diag -r 1)

DCGM (Data Center GPU Manager) diagnostics run in levels. Higher levels take longer and test more.

What NVIDIA says (1)

“Level 1 tests to use as a readiness metric Level 2 tests to use as an epilogue on failure Level 3 and Level 4 tests to be run by an administrator as post-mortem”

— DCGM Diagnostics

Practice 3.5 (4 questions) Full Workload Management guide

← 3.4 Training with Run:ai · 3.6 Sharing resources between teams →