3.5 Human feedback data (RLHF)
How human rankings are collected and used to align models.
Key points
RLHF (reinforcement learning from human feedback) trains a reward model in stage 2. It uses prompts with multiple responses ranked by humans, so the reward model learns to predict human preference.
What NVIDIA says (2)
“A dataset consisting of prompts with multiple responses ranked by humans is used to train the RM to predict human preference.”
“The SFT model is trained as a reward model (RM) in stage 2 of RLHF.”
NVIDIA describes reinforcement learning from human feedback (RLHF) stage 3 as fine-tuning the initial policy model against the reward model using reinforcement learning with proximal policy optimization (PPO).
What NVIDIA says (1)
“stage 3 of RLHF focuses on fine-tuning the initial policy model against the RM using reinforcement learning with a proximal policy optimization (PPO) algorithm.”
Key terms
- Reinforcement learning from human feedback: Aligning an LLM with human preferences using a reward model trained on human rankings.
- Reward model: A model trained on human-ranked responses that predicts which answer people would prefer.
- Proximal policy optimization: The reinforcement-learning algorithm used in RLHF stage 3 to tune the model against the reward model.
Sample question
In RLHF (reinforcement learning from human feedback), what data do humans provide to train the reward model?
Show the answer
Answer: Prompts with multiple responses ranked by humans
RLHF (reinforcement learning from human feedback) trains a reward model in stage 2. It uses prompts with multiple responses ranked by humans, so the reward model learns to predict human preference.
What NVIDIA says (2)
“A dataset consisting of prompts with multiple responses ranked by humans is used to train the RM to predict human preference.”
“The SFT model is trained as a reward model (RM) in stage 2 of RLHF.”
Practice 3.5 (2 questions) Full Experimentation guide
← 3.4 Evaluating technologies · 3.6 Evaluating and benchmarking models →