JA EN
Textbook › Part II The Lineage of AI Models
CHAPTER 26

The Reality of Running at the Edge

★ MEMBER2 min

For the final chapter, we connect Part I and Part II. The subject is how differences in model structure show up on the hardware.

Model form and hardware fit
CNN Transformer / VLM
Dominant operation Convolution (local MACs) Huge matrix products, softmax, normalization
Data reuse High. Weight sharing pays off Low. During generation especially, the weights are re-read every time
Bottleneck Leans toward compute Leans toward memory bandwidth
Memory use Essentially fixed Grows with input length (KV cache)
Support on existing NPUs Almost certainly supported Depends on generation and SDK. Check before committing

The KV cache, a new kind of problem

When a Transformer generates text, every token it emits refers back to the Keys and Values of every token before it. Recomputing them each time would be wasteful, so they are kept around. This is the KV cache.

§

Members-only from here

All 26 chapters and every lab, $4.99/mo. Cancel anytime.

Comments

Sign in to comment