3.5 Monitoring data collection and experiments
Experiment loggers, resuming runs, MLOps tracking and data flywheels.
Key points
Experiment tracking records metrics, settings and checkpoints for each run. NeMo's exp_manager sets this up through PyTorch Lightning.
What NVIDIA says (1)
“The NeMo Framework Experiment Manager leverages PyTorch Lightning for model checkpointing, TensorBoard Logging, Weights and Biases, DLLogger and MLFlow logging.”
A checkpoint is a saved copy of model and optimizer state. With resume_if_exists, the run picks up from the latest checkpoint.
What NVIDIA says (1)
“resume training if checkpoints already exist resume_if_exists : True”
MLOps (machine learning operations) is a set of practices for running AI reliably. Tracking records which data and settings produced each model.
What NVIDIA says (1)
“AI models require careful tracking through cycles of experiments, tuning and retraining.”
Monitoring what users do with a model creates new data. Feeding that data back improves the model over time.
What NVIDIA says (1)
“An AI data flywheel is a self-improving loop where data collected from AI interactions or processes is used to continuously refine AI models”
Key terms
- Experiment Manager: The NeMo component that sets up logging and checkpoints for each training run.
- Checkpoint: A saved copy of training state so a run can resume.
- MLOps: Practices for building, deploying and tracking AI models reliably.
- Data flywheel: A loop where data from AI use improves the model, which then produces better data.
Sample question
Which loggers can the NeMo Experiment Manager write to?
Show the answer
Answer: TensorBoard, Weights and Biases, DLLogger and MLFlow
Experiment tracking records metrics, settings and checkpoints for each run. NeMo's exp_manager sets this up through PyTorch Lightning.
What NVIDIA says (1)
“The NeMo Framework Experiment Manager leverages PyTorch Lightning for model checkpointing, TensorBoard Logging, Weights and Biases, DLLogger and MLFlow logging.”
Practice 3.5 (4 questions) Full Multimodal Data guide
← 3.4 Identifying data, hardware and software components · 3.6 Python packages for traditional ML →