JA EN

#muon

1 articles

01 ·Large Language Models·★ MEMBER·PAPER·13 min read Paper walkthrough: Puro-2B — pretraining a 2B model from scratch for $6.9K on consumer GPUs A team ran 1.4 trillion tokens of pretraining on gaming GPUs and reached Qwen2-1.5B-level quality for roughly $4.4K. Here is the cost structure, the FP8 accounting, the effective learning rate, and the curriculum averaging — from first principles.