2.10 TensorFlow and PyTorch
Automatic mixed precision, the NeMo training stack and data loading in PyTorch.
Key points
Automatic mixed precision lets the framework pick FP16 or FP32 per operation. In supported frameworks it can take one line of code or one environment variable. FP16 means 16-bit floating point. FP32 means 32-bit floating point. NLP means natural-language processing. ML means machine learning.
What NVIDIA says (1)
“Currently, the frameworks with support for automatic mixed precision are TensorFlow, PyTorch, and MXNet.”
NeMo is a PyTorch-based framework. PyTorch Lightning runs the training loop. Hydra manages the YAML configs. YAML means a plain-text configuration file format.
What NVIDIA says (1)
“NeMo uses Hydra for configuring both NeMo models and the PyTorch Lightning Trainer.”
In PyTorch you can swap the data loader for DALI. The example enables it with a --data-backends flag. DALI means the NVIDIA Data Loading Library.
What NVIDIA says (1)
“DALI can use CPU or GPU, and outperforms the PyTorch native dataloader.”
Key terms
- NVIDIA DALI: A GPU-accelerated library for loading and preprocessing image, video and audio data.
- PyTorch Lightning: A PyTorch library that runs the training loop; NeMo builds on it.
- Hydra: A configuration tool that merges YAML files and command-line overrides; NeMo uses it.
- Automatic mixed precision: A framework feature that picks 16-bit or 32-bit math per operation automatically.
Sample question
Which frameworks does NVIDIA list as having automatic mixed precision (AMP) support?
Show the answer
Answer: TensorFlow, PyTorch and MXNet
Automatic mixed precision lets the framework pick FP16 or FP32 per operation. In supported frameworks it can take one line of code or one environment variable. FP16 means 16-bit floating point. FP32 means 32-bit floating point. NLP means natural-language processing. ML means machine learning.
What NVIDIA says (1)
“Currently, the frameworks with support for automatic mixed precision are TensorFlow, PyTorch, and MXNet.”
Practice 2.10 (3 questions) Full Core Machine Learning and AI Knowledge guide
← 2.9 Prompt engineering principles · 3.1 Scalability, performance and reliability →