6.6 Benchmarking and choosing hardware

NCA-ADS · Introductory MLOps Practices (10% of the exam) · Official objective: “Benchmarking workflows and selecting optimal hardware”

Timing CPU and GPU runs fairly, and Multi-Instance GPU for small jobs.

Key points

  1. Benchmarking means timing the same workload on different hardware. Wall-clock time is the real elapsed time. HPO means hyperparameter optimization.

    What NVIDIA says (1)

    “In the notebook demo below, we compare benchmarking results to show how GPU can accelerate HPO tuning jobs relative to CPU.”

    — RAPIDS Deployment: XGBoost and Random Forest GPU vs CPU benchmark

  2. MIG partitions one GPU into isolated instances, each with its own resources. It suits many small workloads; trade-offs include no NVLink between instances. NVLink means NVIDIA's high-speed GPU interconnect.

    What NVIDIA says (1)

    “Multi-Instance GPU is a technology that allows partitioning a single GPU into multiple instances, making each one seem as a completely independent GPU.”

    — RAPIDS Deployment: Multi-Instance GPU (MIG)

Key terms

Sample question

You want to decide whether GPUs are worth it for your HPO workload. What does NVIDIA's benchmark notebook do?

Show the answer

Answer: Runs the same HPO job on GPU and CPU instances and compares wall-clock time

Benchmarking means timing the same workload on different hardware. Wall-clock time is the real elapsed time. HPO means hyperparameter optimization.

What NVIDIA says (1)

“In the notebook demo below, we compare benchmarking results to show how GPU can accelerate HPO tuning jobs relative to CPU.”

— RAPIDS Deployment: XGBoost and Random Forest GPU vs CPU benchmark

Practice 6.6 (2 questions) Full Introductory MLOps Practices guide

← 6.5 Artifacts and configurations for reproducibility · 7.1 Time series splits and forecast evaluation →