Your next step
Learning from feedback
Preferences, reward models, RLHF, DPO, and reward gaming.
Start learningWhat you’ll learn
People can compare two responses
Identify preference data rather than a single correct answer.
About 4 minutes · Open activity
A reward model predicts preference
Distinguish a learned score from a verified truth judgment.
About 4 minutes · Open activity
RLHF uses human feedback to guide training
Trace preference evidence into a further-training signal.
About 4 minutes · Open activity
DPO learns directly from preference pairs
Recognize a different route for preference-based adaptation.
About 4 minutes · Open activity
AI-generated feedback still has human choices behind it
Explain what shapes RLAIF or constitutional feedback.
About 4 minutes · Open activity
Some outcomes can be checked automatically
Recognize verifiable rewards on a bounded task.
About 4 minutes · Open activity
Optimizing the score can miss the intention
Identify reward gaming in a novel scenario.
About 4 minutes · Open activity
Safer behavior needs continuing evaluation
Distinguish trained refusal behavior from guaranteed safety.
About 4 minutes · Open activity
Improve the feedback rubric
Compare two response preferences, identify a reward shortcut, and propose a bounded check that does not confuse approval with truth.
About 6 minutes · Open activity