JA EN
Glossary › rewardbench

GLOSSARY

rewardbench

appears in 1 paper titles

Definition

A benchmark for reward models, released by the Allen Institute for AI. It supplies pairs of chosen and rejected responses and checks whether the reward model scores the preferred one higher, broken out by category — chat, reasoning, safety, and adversarial cases. What it grades is not a generator but the scorer that RLHF and preference tuning are built on, which is why a weak result here propagates into every model aligned with that reward model.