6.1 Efficiency versus accuracy
Tensor Core alignment, memory-bound layers, BF16 and activation recomputation.
Key points
Tensor Cores are GPU units for fast matrix math. They work best when sizes line up with their tile shapes: multiples of 4 for TF32, 8 for FP16 and 16 for INT8. FP16 means 16-bit floating point. TF32 means TensorFloat-32, a 19-bit Tensor Core math format. INT8 means 8-bit integer.
What NVIDIA says (1)
“Tensor Cores are most efficient when key parameters of the operation are multiples of 4 if using TF32, 8 if using FP16, or 16 if using INT8”
Math-limited means compute speed is the bottleneck. Memory-bound means moving data is the bottleneck. Small layer parameters often make an operation memory-bound.
What NVIDIA says (1)
“speeding up calculation does not improve performance.”
BF16 is a 16-bit format with the same exponent range as FP32. It saves memory and time while staying numerically stable. CLIP means Contrastive Language-Image Pre-training. FP16 means 16-bit floating point. FP32 means 32-bit floating point.
What NVIDIA says (1)
“Training is conducted in Bfloat16 precision, which offers a balance between the higher precision of FP32 and the memory savings and speed of FP16.”
Activation recomputation stores only some activations and recomputes the rest in the backward pass. It trades extra compute for less memory.
What NVIDIA says (2)
“Checkpointing a few activations and recomputing the rest is a common technique to reduce device memory usage.”
“it increases the per-transformer layer computation cost by 30%”
Key terms
- Mixed precision: Training mostly in 16-bit formats while keeping key values in 32-bit, to save memory and time.
- bfloat16: A 16-bit number format with FP32's range, used to train faster with less memory.
- Tensor Core: A GPU unit for fast matrix math that works best when layer sizes are aligned.
- Memory-bound: An operation limited by data movement rather than math speed.
- Activation recomputation: Storing only some activations and recomputing the rest in the backward pass to save memory.
Sample question
For FP16 Tensor Core math, which layer sizes run most efficiently?
Show the answer
Answer: Key parameters that are multiples of 8
Tensor Cores are GPU units for fast matrix math. They work best when sizes line up with their tile shapes: multiples of 4 for TF32, 8 for FP16 and 16 for INT8. FP16 means 16-bit floating point. TF32 means TensorFloat-32, a 19-bit Tensor Core math format. INT8 means 8-bit integer.
What NVIDIA says (1)
“Tensor Cores are most efficient when key parameters of the operation are multiples of 4 if using TF32, 8 if using FP16, or 16 if using INT8”
Practice 6.1 (4 questions) Full Performance Optimization guide
← 5.4 Relationships, trends and confounders · 6.2 Hyperparameter tuning →