Glossary › preference
GLOSSARY
preference
appears in 2 paper titles
Definition
A comparison judgment: given two or more candidate outputs, which one a human (or a model acting as judge) prefers. Pairwise preferences are easier to collect and more consistent across annotators than absolute scores, which is why RLHF and DPO train on them. The limitation is that a preference records the choice but not the reason, so systematic annotator biases — toward length, confidence, or formatting — get optimized into the model along with genuine quality.
Explainers using this term
- DPO and What Came After — The Lineage That Simplified RLHFDirect Preference Optimization: Your Language Model is Secretly a Reward Model