2.1 Training stability

NCA-GENM · Core Machine Learning and AI Knowledge (20% of the exam) · Official objective: “Control training stability in multimodal settings.”

Convergence settings in CLIP and loss scaling in mixed precision.

Key points

  1. Convergence means the loss settles to a good value as training goes on. gather_with_grad keeps full gradients when features are gathered across GPUs. Turning it off can make training unstable. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “Disabling this (setting to False ) may cause convergence issues.”

    — NeMo Framework 24.09: CLIP

  2. Mixed precision trains mostly in 16-bit floating point (FP16) for speed. Small gradients can round to zero in FP16. Loss scaling multiplies the loss so those gradients stay representable. FP32 means 32-bit floating point. INT8 means 8-bit integer.

    What NVIDIA says (1)

    “Mixed precision without loss scaling (grey) diverges after a while, whereas mixed precision with loss scaling (green) matches the single precision model (black).”

    — Train With Mixed Precision

  3. An overflow means a value became too large (Inf or NaN). Dynamic loss scaling then skips that step and lowers the scale. If no overflow happens for a while, it raises the scale again.

    What NVIDIA says (1)

    “If an overflow occurs, skip the weight update and decrease the scaling factor.”

    — Train With Mixed Precision

Key terms

Sample question

In NeMo CLIP training on many GPUs, what can happen if you set gather_with_grad to False?

Show the answer

Answer: Training may have convergence issues

Convergence means the loss settles to a good value as training goes on. gather_with_grad keeps full gradients when features are gathered across GPUs. Turning it off can make training unstable. CLIP means Contrastive Language-Image Pre-training.

What NVIDIA says (1)

“Disabling this (setting to False ) may cause convergence issues.”

— NeMo Framework 24.09: CLIP

Practice 2.1 (3 questions) Full Core Machine Learning and AI Knowledge guide

← 1.5 Testing models for accuracy · 2.2 Multimodal loss functions →