PaperLens
紙
Students
Professional
JA
EN
◐
Sign in with Google
Sign in
Read
Home
Close reading
New
Textbook
Go deeper
Learn
Lab
Landscape
Contributors
Glossary
You
Search
All-access
My Page
#data-efficiency
1 articles
01
2026-09-07
·
Inference & Serving
·
★ MEMBER
·
PAPER
·
11 min read
Paper Walkthrough: One Training Example Keeps On-Policy Distillation Improving for Hundreds of Steps
Trained on a single query, on-policy distillation still improves for hundreds of steps and recovers most of full-data OPD's gain. The paper explains this with state coverage and absorption rate, and concludes OPD is data-overfed but algorithm-starved.