JA EN

#common-crawl

1 articles

01 ·Large Language Models·★ MEMBER·PAPER·10 min read Building a Pretraining Corpus — From Web Sludge to Textbook Quality Behind the single line "pretrained on a large corpus of web text" sit four stages: text extraction, quality filtering, deduplication, and mixture weights. This is a from-scratch walkthrough of how a gravel heap called Common Crawl gets sifted into textbook-quality prose — from the MinHash equation to the parameter names you actually touch.