Glossary › reward
GLOSSARY
reward
appears in 5 paper titles
Definition
The scalar signal in reinforcement learning that tells the agent how good an action or outcome was; the policy is updated to increase expected reward. For language models the reward often comes from a learned reward model trained on human preference comparisons, standing in for expensive human judgement. Because policies optimize exactly what is measured, poorly specified rewards invite reward hacking — high scores achieved by behaviour nobody wanted.
Explainers using this term
- DPO and What Came After — The Lineage That Simplified RLHFDirect Preference Optimization: Your Language Model is Secretly a Reward Model
- Paper Walkthrough: SA-MRPO — Stop Studying the Subject You've Already AcedLearn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
- Paper Explained: OSReward — Can You Trust the AI That Grades AI? Remeasuring Rewards for Computer-Use AgentsOSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models