JA EN

#alphazero

1 articles

01 ·Evaluation & Judging·★ MEMBER·PAPER·12 min read Self-Improving AI — Self-Play, Co-Evolution, and Generated Curricula What has to be true for a model to get better without anyone adding data? This article pulls three conditions out of AlphaZero's self-play, shows exactly which one breaks first for language models, explains how co-evolution and generated curricula try to patch the gap, and ends with why self-improvement claims are unusually easy to evaluate wrong.