6.2 Hyperparameter tuning
Optimizer and schedule configs, Hydra overrides, LoRA rank and batch sizing.
Key points
Learning rate and warmup are hyperparameters, set before training. Warmup ramps the learning rate up gradually at the start.
What NVIDIA says (1)
“Optimizers and learning rate schedules are configurable across all NeMo models and have their own namespace.”
Hydra merges settings from several layers. Command-line overrides are handy for quick tuning sweeps. YAML means a plain-text configuration file format.
What NVIDIA says (1)
“Configuration with Hydra always has the following precedence CLI > YAML > Dataclass.”
LoRA (Low-Rank Adaptation) trains small low-rank matrices while freezing the base model. The rank r is a hyperparameter you tune.
What NVIDIA says (2)
“is a hyperparameter that controls the rank of the decomposition”
“Choosing a smaller \(r\) can save a lot of parameters and memory and achieve faster training. However, a smaller \(r\) can potentially decrease task-specific information captu”
GPUs split work into tiles. Sizes that divide evenly avoid wasted work.
What NVIDIA says (1)
“choosing parameters (including batch size, input size, output size, and channel counts) to be divisible by larger powers of two, at least 64, and up to 256.”
Key terms
- Low-Rank Adaptation: A PEFT method that trains small low-rank matrices; its rank r is a hyperparameter.
- Hydra: A configuration tool that merges YAML files and command-line overrides; NeMo uses it.
- Hyperparameter: A setting chosen before training, such as learning rate or LoRA rank.
Try it
Sample question
Where do you set the optimizer and learning-rate schedule for a NeMo model?
Show the answer
Answer: In the model's optim config namespace, for example with warmup_steps
Learning rate and warmup are hyperparameters, set before training. Warmup ramps the learning rate up gradually at the start.
What NVIDIA says (1)
“Optimizers and learning rate schedules are configurable across all NeMo models and have their own namespace.”
Practice 6.2 (4 questions) Full Performance Optimization guide
← 6.1 Efficiency versus accuracy · 6.3 Transfer learning for performance →