Parallel & Distributed
Limits of parallelism, the GPU execution model, communication in distributed training
01
·Parallel & Distributed·★ MEMBER·10 min read
Why GPUs Are Fast — The Execution Model and the Limits of Parallelism
CPUs and GPUs do not mean the same thing by fast. Where the transistor budget goes, how SIMT bundles 32 threads into a warp, why branch divergence costs you, occupancy and register pressure — and finally Amdahl's law as a way to bound the payoff before you start, plus the profiler counters that tell you when the CPU is the bottleneck.
02
·Parallel & Distributed·★ MEMBER·10 min read
Distributed Training from Scratch — Data Parallel, Model Parallel, and When Communication Becomes the Bottleneck
Why one machine is not enough, counted out in bytes; data parallelism and all-reduce; what ZeRO and FSDP actually shard; tensor and pipeline parallelism. Then the ratio of computation to communication that tells you where scaling stops paying — and gradient accumulation, NCCL settings and how to diagnose a hang.
03
·Parallel & Distributed·★ MEMBER·9 min read
Concurrency from Scratch — Locks, Atomics, and Memory Models
Why data races happen and why they refuse to reproduce in your tests, starting from zero. Locks, atomic operations, CAS, and memory models, built up through metaphor, math, and code.