JA EN

#flops

1 articles

01 ·Numerical Computing·★ MEMBER·8 min read The Cost of Matrix Multiplication — Where Almost All of AI's Compute Goes Why GEMM is everything: the anatomy of O(n³), memory bandwidth and arithmetic intensity, what actually makes a GPU fast, and the intuition behind tiling — ending with you able to estimate a model's training and inference FLOPs yourself.