JA EN

#preference-learning

1 articles

01 ·Training & Alignment·★ MEMBER·PAPER·13 min read DPO and What Came After — The Lineage That Simplified RLHF Derives DPO one line at a time, starting from the closed-form solution to KL-constrained reward maximization, to show why no separate reward model is needed. Then organizes IPO (which explains DPO's overfitting mathematically), KTO (which drops the pairing requirement), and GRPO (which drops the value model and goes back online) by what each one deleted — and gives a rule for choosing based on the shape of the data you actually have.