Gen AI Lingo

Your next step

Learning from feedback

Preferences, reward models, RLHF, DPO, and reward gaming.

Start learning

What you’ll learn

  1. People can compare two responses

    Identify preference data rather than a single correct answer.

    About 4 minutes · Open activity

  2. A reward model predicts preference

    Distinguish a learned score from a verified truth judgment.

    About 4 minutes · Open activity

  3. RLHF uses human feedback to guide training

    Trace preference evidence into a further-training signal.

    About 4 minutes · Open activity

  4. DPO learns directly from preference pairs

    Recognize a different route for preference-based adaptation.

    About 4 minutes · Open activity

  5. AI-generated feedback still has human choices behind it

    Explain what shapes RLAIF or constitutional feedback.

    About 4 minutes · Open activity

  6. Some outcomes can be checked automatically

    Recognize verifiable rewards on a bounded task.

    About 4 minutes · Open activity

  7. Optimizing the score can miss the intention

    Identify reward gaming in a novel scenario.

    About 4 minutes · Open activity

  8. Safer behavior needs continuing evaluation

    Distinguish trained refusal behavior from guaranteed safety.

    About 4 minutes · Open activity

  9. Improve the feedback rubric

    Compare two response preferences, identify a reward shortcut, and propose a bounded check that does not confuse approval with truth.

    About 6 minutes · Open activity