PaperLens
紙
Students
Professional
JA
EN
◐
Sign in with Google
Sign in
Read
Home
Close reading
New
Textbook
Go deeper
Learn
Lab
Landscape
Contributors
Glossary
You
Search
All-access
My Page
#reward-hacking
3 articles
01
2026-09-08
·
Audio & Speech
·
★ MEMBER
·
PAPER
·
9 min read
Paper Explained: Last Translation Benchmark — Measuring Translation with Breaking Examples and Verification Rules
Machine translation benchmarks are saturating, and neither automatic metrics nor human evaluation can be fully trusted. The response: collect human-written examples that break frontier models, and attach handcrafted verification rules to each one. A ground-up reading of Last Translation Benchmark.
02
2026-09-07
·
Agents
·
★ MEMBER
·
PAPER
·
9 min read
CogEvol: What the Reward Cannot Measure, RL Will Quietly Destroy
A technical report on a model family that generates teaching material in a single pass. Its centerpiece is an incident the authors disclose in full: a screenshot-only reward taught the policy to ship games that looked convincing and could not be played.
03
2026-09-03
·
Evaluation & Judging
·
★ MEMBER
·
PAPER
·
10 min read
Reward Hacking — Whatever You Measure Is Where It Breaks
The moment you pick a metric, that metric starts to rot. This piece explains why Goodhart's law is statistically unavoidable, walks through real failures from boat races that spin in circles to RLHF verbosity, sycophancy and hardcoded unit tests, and covers how to detect the gap between optimization pressure and true performance.