Glossary › rewardbench
GLOSSARY
rewardbench
appears in 1 paper titles
Definition
A benchmark for reward models, released by the Allen Institute for AI. It supplies pairs of chosen and rejected responses and checks whether the reward model scores the preferred one higher, broken out by category — chat, reasoning, safety, and adversarial cases. What it grades is not a generator but the scorer that RLHF and preference tuning are built on, which is why a weak result here propagates into every model aligned with that reward model.