Glossary › compute-optimal
GLOSSARY
compute-optimal
appears in 3 paper titles
Definition
Allocating a fixed training budget between model size and number of training tokens so as to minimise loss. The Chinchilla result popularised the idea and showed that the large models of its day were undertrained: too many parameters for their data. Note that the objective ignores inference: if a model will be served heavily, it is rational to go smaller and train past the compute-optimal point: the extra training cost is paid once, the serving saving recurs.
Explainers using this term
- Scaling Skepticism — A Genealogy of the "Just Make It Bigger" CritiqueTraining Compute-Optimal Large Language Models
- Scaling Laws from Scratch — Why Making Models Bigger Makes Them Smarter (and When It Doesn't)Training Compute-Optimal Large Language Models