JA EN

#sparse-model

1 articles

01 ·Paper Deep-Dives·★ MEMBER·PAPER·9 min read Mixture of Experts (MoE) from Scratch — Routing and Load Balancing in the Switch Transformer A ground-up explanation of Mixture of Experts, the sparse architecture behind today's largest LLMs, built strictly from the Switch Transformer paper: the router math, expert capacity, the load-balancing loss, and the three tricks that make sparse training stable.