JA EN

#label-free

1 articles

01 ·Agents·★ MEMBER·PAPER·12 min read Paper Explained: Co-RL — Reasoning Without Labels, Emerging From a Diverse Cohort Grade your own answers long enough and the model collapses. Co-RL breaks that loop by rewarding each agent against a peer's majority vote, matching supervised training without touching a single ground-truth label. The mechanism, the dynamics, the numbers, and the traps — straight from the paper.