JA EN

#fsdp

1 articles

01 ·Parallel & Distributed·★ MEMBER·10 min read Distributed Training from Scratch — Data Parallel, Model Parallel, and When Communication Becomes the Bottleneck Why one machine is not enough, counted out in bytes; data parallelism and all-reduce; what ZeRO and FSDP actually shard; tensor and pipeline parallelism. Then the ratio of computation to communication that tells you where scaling stops paying — and gradient accumulation, NCCL settings and how to diagnose a hang.