3.5 Human feedback data (RLHF)

NCA-GENL · Experimentation (22% of the exam) · Official objective: “Awareness of, and/or participation in, data collection from human subjects (e.g., RLHF).”

How human rankings are collected and used to align models.

Key points

  1. RLHF (reinforcement learning from human feedback) trains a reward model in stage 2. It uses prompts with multiple responses ranked by humans, so the reward model learns to predict human preference.

    What NVIDIA says (2)

    “A dataset consisting of prompts with multiple responses ranked by humans is used to train the RM to predict human preference.”

    — Mastering LLM Techniques: Customization

    “The SFT model is trained as a reward model (RM) in stage 2 of RLHF.”

    — Mastering LLM Techniques: Customization

  2. NVIDIA describes reinforcement learning from human feedback (RLHF) stage 3 as fine-tuning the initial policy model against the reward model using reinforcement learning with proximal policy optimization (PPO).

    What NVIDIA says (1)

    “stage 3 of RLHF focuses on fine-tuning the initial policy model against the RM using reinforcement learning with a proximal policy optimization (PPO) algorithm.”

    — Mastering LLM Techniques: Customization

Key terms

Sample question

In RLHF (reinforcement learning from human feedback), what data do humans provide to train the reward model?

Show the answer

Answer: Prompts with multiple responses ranked by humans

RLHF (reinforcement learning from human feedback) trains a reward model in stage 2. It uses prompts with multiple responses ranked by humans, so the reward model learns to predict human preference.

What NVIDIA says (2)

“A dataset consisting of prompts with multiple responses ranked by humans is used to train the RM to predict human preference.”

— Mastering LLM Techniques: Customization

“The SFT model is trained as a reward model (RM) in stage 2 of RLHF.”

— Mastering LLM Techniques: Customization

Practice 3.5 (2 questions) Full Experimentation guide

← 3.4 Evaluating technologies · 3.6 Evaluating and benchmarking models →