JA EN

#occupancy

1 articles

01 ·Parallel & Distributed·★ MEMBER·10 min read Why GPUs Are Fast — The Execution Model and the Limits of Parallelism CPUs and GPUs do not mean the same thing by fast. Where the transistor budget goes, how SIMT bundles 32 threads into a warp, why branch divergence costs you, occupancy and register pressure — and finally Amdahl's law as a way to bound the payoff before you start, plus the profiler counters that tell you when the CPU is the bottleneck.