2.1 Training stability
Convergence settings in CLIP and loss scaling in mixed precision.
Key points
Convergence means the loss settles to a good value as training goes on. gather_with_grad keeps full gradients when features are gathered across GPUs. Turning it off can make training unstable. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“Disabling this (setting to False ) may cause convergence issues.”
Mixed precision trains mostly in 16-bit floating point (FP16) for speed. Small gradients can round to zero in FP16. Loss scaling multiplies the loss so those gradients stay representable. FP32 means 32-bit floating point. INT8 means 8-bit integer.
What NVIDIA says (1)
“Mixed precision without loss scaling (grey) diverges after a while, whereas mixed precision with loss scaling (green) matches the single precision model (black).”
An overflow means a value became too large (Inf or NaN). Dynamic loss scaling then skips that step and lowers the scale. If no overflow happens for a while, it raises the scale again.
What NVIDIA says (1)
“If an overflow occurs, skip the weight update and decrease the scaling factor.”
Key terms
- Loss scaling: Multiplying the loss before backpropagation so small FP16 gradients do not round to zero.
- Mixed precision: Training mostly in 16-bit formats while keeping key values in 32-bit, to save memory and time.
Sample question
In NeMo CLIP training on many GPUs, what can happen if you set gather_with_grad to False?
Show the answer
Answer: Training may have convergence issues
Convergence means the loss settles to a good value as training goes on. gather_with_grad keeps full gradients when features are gathered across GPUs. Turning it off can make training unstable. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“Disabling this (setting to False ) may cause convergence issues.”
Practice 2.1 (3 questions) Full Core Machine Learning and AI Knowledge guide
← 1.5 Testing models for accuracy · 2.2 Multimodal loss functions →