3.5 Monitoring data collection and experiments

NCA-GENM · Multimodal Data (15% of the exam) · Official objective: “Monitor the functioning of data collection, experiments, and other software processes.”

Experiment loggers, resuming runs, MLOps tracking and data flywheels.

Key points

  1. Experiment tracking records metrics, settings and checkpoints for each run. NeMo's exp_manager sets this up through PyTorch Lightning.

    What NVIDIA says (1)

    “The NeMo Framework Experiment Manager leverages PyTorch Lightning for model checkpointing, TensorBoard Logging, Weights and Biases, DLLogger and MLFlow logging.”

    — NeMo Framework 24.09: Experiment Manager

  2. A checkpoint is a saved copy of model and optimizer state. With resume_if_exists, the run picks up from the latest checkpoint.

    What NVIDIA says (1)

    “resume training if checkpoints already exist resume_if_exists : True”

    — NeMo Framework 24.09: Experiment Manager

  3. MLOps (machine learning operations) is a set of practices for running AI reliably. Tracking records which data and settings produced each model.

    What NVIDIA says (1)

    “AI models require careful tracking through cycles of experiments, tuning and retraining.”

    — What is MLOps?

  4. Monitoring what users do with a model creates new data. Feeding that data back improves the model over time.

    What NVIDIA says (1)

    “An AI data flywheel is a self-improving loop where data collected from AI interactions or processes is used to continuously refine AI models”

    — Data flywheel: What it is and how it works

Key terms

Sample question

Which loggers can the NeMo Experiment Manager write to?

Show the answer

Answer: TensorBoard, Weights and Biases, DLLogger and MLFlow

Experiment tracking records metrics, settings and checkpoints for each run. NeMo's exp_manager sets this up through PyTorch Lightning.

What NVIDIA says (1)

“The NeMo Framework Experiment Manager leverages PyTorch Lightning for model checkpointing, TensorBoard Logging, Weights and Biases, DLLogger and MLFlow logging.”

— NeMo Framework 24.09: Experiment Manager

Practice 3.5 (4 questions) Full Multimodal Data guide

← 3.4 Identifying data, hardware and software components · 3.6 Python packages for traditional ML →