6.2 Hyperparameter tuning

NCA-GENM · Performance Optimization (10% of the exam) · Official objective: “Optimize the performance of AI models, including tuning hyperparameters.”

Optimizer and schedule configs, Hydra overrides, LoRA rank and batch sizing.

Key points

  1. Learning rate and warmup are hyperparameters, set before training. Warmup ramps the learning rate up gradually at the start.

    What NVIDIA says (1)

    “Optimizers and learning rate schedules are configurable across all NeMo models and have their own namespace.”

    — NeMo Framework 24.09: NeMo Models (core)

  2. Hydra merges settings from several layers. Command-line overrides are handy for quick tuning sweeps. YAML means a plain-text configuration file format.

    What NVIDIA says (1)

    “Configuration with Hydra always has the following precedence CLI > YAML > Dataclass.”

    — NeMo Framework 24.09: NeMo Models (core)

  3. LoRA (Low-Rank Adaptation) trains small low-rank matrices while freezing the base model. The rank r is a hyperparameter you tune.

    What NVIDIA says (2)

    “is a hyperparameter that controls the rank of the decomposition”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

    “Choosing a smaller \(r\) can save a lot of parameters and memory and achieve faster training. However, a smaller \(r\) can potentially decrease task-specific information captu”

    — Tune and Deploy LoRA LLMs with NVIDIA TensorRT-LLM

  4. GPUs split work into tiles. Sizes that divide evenly avoid wasted work.

    What NVIDIA says (1)

    “choosing parameters (including batch size, input size, output size, and channel counts) to be divisible by larger powers of two, at least 64, and up to 256.”

    — NVIDIA Deep Learning Performance: Getting Started

Key terms

Try it

Sample question

Where do you set the optimizer and learning-rate schedule for a NeMo model?

Show the answer

Answer: In the model's optim config namespace, for example with warmup_steps

Learning rate and warmup are hyperparameters, set before training. Warmup ramps the learning rate up gradually at the start.

What NVIDIA says (1)

“Optimizers and learning rate schedules are configurable across all NeMo models and have their own namespace.”

— NeMo Framework 24.09: NeMo Models (core)

Practice 6.2 (4 questions) Full Performance Optimization guide

← 6.1 Efficiency versus accuracy · 6.3 Transfer learning for performance →