2.10 TensorFlow and PyTorch

NCA-GENM · Core Machine Learning and AI Knowledge (20% of the exam) · Official objective: “Work with deep learning frameworks such as TensorFlow or PyTorch.”

Automatic mixed precision, the NeMo training stack and data loading in PyTorch.

Key points

  1. Automatic mixed precision lets the framework pick FP16 or FP32 per operation. In supported frameworks it can take one line of code or one environment variable. FP16 means 16-bit floating point. FP32 means 32-bit floating point. NLP means natural-language processing. ML means machine learning.

    What NVIDIA says (1)

    “Currently, the frameworks with support for automatic mixed precision are TensorFlow, PyTorch, and MXNet.”

    — Train With Mixed Precision

  2. NeMo is a PyTorch-based framework. PyTorch Lightning runs the training loop. Hydra manages the YAML configs. YAML means a plain-text configuration file format.

    What NVIDIA says (1)

    “NeMo uses Hydra for configuring both NeMo models and the PyTorch Lightning Trainer.”

    — NeMo Framework 24.09: NeMo Models (core)

  3. In PyTorch you can swap the data loader for DALI. The example enables it with a --data-backends flag. DALI means the NVIDIA Data Loading Library.

    What NVIDIA says (1)

    “DALI can use CPU or GPU, and outperforms the PyTorch native dataloader.”

    — NVIDIA DeepLearningExamples: ResNet50 v1.5 for PyTorch

Key terms

Sample question

Which frameworks does NVIDIA list as having automatic mixed precision (AMP) support?

Show the answer

Answer: TensorFlow, PyTorch and MXNet

Automatic mixed precision lets the framework pick FP16 or FP32 per operation. In supported frameworks it can take one line of code or one environment variable. FP16 means 16-bit floating point. FP32 means 32-bit floating point. NLP means natural-language processing. ML means machine learning.

What NVIDIA says (1)

“Currently, the frameworks with support for automatic mixed precision are TensorFlow, PyTorch, and MXNet.”

— Train With Mixed Precision

Practice 2.10 (3 questions) Full Core Machine Learning and AI Knowledge guide

← 2.9 Prompt engineering principles · 3.1 Scalability, performance and reliability →