6.1 Efficiency versus accuracy

NCA-GENM · Performance Optimization (10% of the exam) · Official objective: “Enhance computational efficiency and improve the accuracy of outputs in AI models.”

Tensor Core alignment, memory-bound layers, BF16 and activation recomputation.

Key points

  1. Tensor Cores are GPU units for fast matrix math. They work best when sizes line up with their tile shapes: multiples of 4 for TF32, 8 for FP16 and 16 for INT8. FP16 means 16-bit floating point. TF32 means TensorFloat-32, a 19-bit Tensor Core math format. INT8 means 8-bit integer.

    What NVIDIA says (1)

    “Tensor Cores are most efficient when key parameters of the operation are multiples of 4 if using TF32, 8 if using FP16, or 16 if using INT8”

    — NVIDIA Deep Learning Performance: Getting Started

  2. Math-limited means compute speed is the bottleneck. Memory-bound means moving data is the bottleneck. Small layer parameters often make an operation memory-bound.

    What NVIDIA says (1)

    “speeding up calculation does not improve performance.”

    — NVIDIA Deep Learning Performance: Getting Started

  3. BF16 is a 16-bit format with the same exponent range as FP32. It saves memory and time while staying numerically stable. CLIP means Contrastive Language-Image Pre-training. FP16 means 16-bit floating point. FP32 means 32-bit floating point.

    What NVIDIA says (1)

    “Training is conducted in Bfloat16 precision, which offers a balance between the higher precision of FP32 and the memory savings and speed of FP16.”

    — NeMo Framework 24.09: CLIP

  4. Activation recomputation stores only some activations and recomputes the rest in the backward pass. It trades extra compute for less memory.

    What NVIDIA says (2)

    “Checkpointing a few activations and recomputing the rest is a common technique to reduce device memory usage.”

    — NeMo Framework 24.09: Activation Recomputation

    “it increases the per-transformer layer computation cost by 30%”

    — NeMo Framework 24.09: Activation Recomputation

Key terms

Sample question

For FP16 Tensor Core math, which layer sizes run most efficiently?

Show the answer

Answer: Key parameters that are multiples of 8

Tensor Cores are GPU units for fast matrix math. They work best when sizes line up with their tile shapes: multiples of 4 for TF32, 8 for FP16 and 16 for INT8. FP16 means 16-bit floating point. TF32 means TensorFloat-32, a 19-bit Tensor Core math format. INT8 means 8-bit integer.

What NVIDIA says (1)

“Tensor Cores are most efficient when key parameters of the operation are multiples of 4 if using TF32, 8 if using FP16, or 16 if using INT8”

— NVIDIA Deep Learning Performance: Getting Started

Practice 6.1 (4 questions) Full Performance Optimization guide

← 5.4 Relationships, trends and confounders · 6.2 Hyperparameter tuning →