6.2 Tracking experiments

NCA-ADS · Introductory MLOps Practices (10% of the exam) · Official objective: “Managing and tracking experiments with MLflow, Weights & Biases, and custom tools”

MLflow, Weights & Biases and why every run must be traceable.

Key points

  1. Experiment tracking records what was run, with which settings, and how it scored. That lets you compare runs and reproduce the best one. ML means machine learning. XGBoost means eXtreme Gradient Boosting.

    What NVIDIA says (1)

    “Experiment tracking: MLflow logs hyperparameters, model artifacts, and evaluation metrics for every run”

    — RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

  2. Weights & Biases (W&B) is an experiment tracking platform. Like MLflow, it records runs so you can debug and reproduce models. ML means machine learning.

    What NVIDIA says (1)

    “Weights & Biases : W&B’s tools help many ML users build better models faster, debugging and reproducing their models with just a few lines of code”

    — What is MLOps?

  3. MLOps means machine learning operations: practices for running AI reliably in production. Tracking ties each model to its data, code and settings.

    What NVIDIA says (1)

    “AI models require careful tracking through cycles of experiments, tuning and retraining.”

    — What is MLOps?

Key terms

Sample question

What does MLflow do in NVIDIA's fraud detection pipeline?

Show the answer

Answer: It logs hyperparameters, model artifacts and evaluation metrics for every run, and manages the model registry

Experiment tracking records what was run, with which settings, and how it scored. That lets you compare runs and reproduce the best one. ML means machine learning. XGBoost means eXtreme Gradient Boosting.

What NVIDIA says (1)

“Experiment tracking: MLflow logs hyperparameters, model artifacts, and evaluation metrics for every run”

— RAPIDS Deployment: Fraud detection model lifecycle with cuDF, Prefect, MLflow and Triton

Practice 6.2 (3 questions) Full Introductory MLOps Practices guide

← 6.1 Monitoring and optimizing ML pipelines · 6.3 Saving, loading and predicting →