4.9 ClusterKit node assessment

NCP-AII · Cluster Test and Verification (33% of the exam) · Official objective: “Run ClusterKit to perform a multifaceted node assessment.”

What ClusterKit tests, how it flags bad results and how scopes and stress runs work.

Key points

  1. ClusterKit is part of HPC-X, NVIDIA's HPC communication toolkit. It tests latency, bandwidth, GPU communication, collectives and CPU/GPU stress. That is why NVIDIA calls it multifaceted.

    What NVIDIA says (3)

    “ClusterKit is a multipurpose node assessment tool for high-performance clusters”

    — HPC-X: ClusterKit

    “SLURM or passwordless ssh connectivity across the hosts.”

    — HPC-X: ClusterKit

    “ClusterKit is the recommended tool for InfiniBand end-to-end performance validation.”

    — InfiniBand Cluster Bring-up Procedure: Performance Testing

  2. A pairwise test measures pairs of nodes. ClusterKit judges each pair against the best result, not a fixed number. Latency over 2.1 times the minimum is also bad by default.

    What NVIDIA says (2)

    “Message bandwidths (BWs) less than 93% (by default) of the maximum”

    — HPC-X: Running ClusterKit

    “Message latencies that are 2.1 times (by default) above the minimum are considered 'bad,'”

    — HPC-X: Running ClusterKit

  3. Oversubscription means a switch has less uplink bandwidth than downlink bandwidth. Cross-rack links then measure lower by design. With --topo-file, ClusterKit divides the expected maximum by the oversubscription ratio.

    What NVIDIA says (2)

    “ClusterKit can account for oversubscription using a topology information file.”

    — HPC-X: Running ClusterKit

    “The ratio of total downlink bandwidth to total uplink bandwidth is referred to as the oversubscription ratio.”

    — HPC-X: Running ClusterKit

  4. A scope is a set of nodes, such as one rack. The scope_info file defines them. Comparing scopes shows whether one rack behaves differently from the others.

    What NVIDIA says (1)

    “The purpose of scoped tests is to analyze how similar sets of nodes (such as all nodes in a single rack) behave and whether there are differences among them.”

    — HPC-X: Running ClusterKit

  5. Testing under stress shows how the fabric behaves when nodes are hot and busy. --with-stress loads CPU, GPU or both. -Y sets the duration in minutes.

    What NVIDIA says (2)

    “To run stress tests, use the --with-stress[=<STRESS_TYPES>] option with the ClusterKit binary.”

    — HPC-X: Running ClusterKit

    “If the tests finish earlier than the specified <TEST_TIME> , they will re-run.”

    — HPC-X: Running ClusterKit

  6. A per-GPU pair test checks each GPU's network path on its own, so one bad GPU-NIC path stands out. -z makes the GPU-to-GPU tests pair GPUs by index.

    What NVIDIA says (1)

    “test corresponding GPU pairs: GPU0-to-GPU0, GPU1-to-GPU1, etc.”

    — HPC-X: ClusterKit Test Descriptions and Options

Key terms

Sample question

What is ClusterKit, and what does it need to run across nodes?

Show the answer

Answer: A multipurpose node assessment tool. It needs Slurm or passwordless SSH across the hosts.

ClusterKit is part of HPC-X, NVIDIA's HPC communication toolkit. It tests latency, bandwidth, GPU communication, collectives and CPU/GPU stress. That is why NVIDIA calls it multifaceted.

What NVIDIA says (3)

“ClusterKit is a multipurpose node assessment tool for high-performance clusters”

— HPC-X: ClusterKit

“SLURM or passwordless ssh connectivity across the hosts.”

— HPC-X: ClusterKit

“ClusterKit is the recommended tool for InfiniBand end-to-end performance validation.”

— InfiniBand Cluster Bring-up Procedure: Performance Testing

Practice 4.9 (6 questions) Full Cluster Test and Verification guide

← 4.8 Transceiver firmware · 4.10 NCCL east-west bandwidth →