Glossary › ttpo
GLOSSARY
ttpo
appears in 1 paper titles
Definition
A paper-coined method name in the growing family of "-PO" objectives, where PO stands for preference optimization. These methods descend from DPO: instead of training a separate reward model and running RL, they update the policy directly from pairs of preferred and rejected outputs. The leading letters vary by paper and encode what that work changes, so the acronym only means what its own paper defines.
Explainers using this term
- TTPO Explained: Training a Model Mid-Exam, With No Answer KeyTTPO: Test-Time Policy Optimization