Textbook › Part II The Lineage of AI Models
CHAPTER 26
The Reality of Running at the Edge
For the final chapter, we connect Part I and Part II. The subject is how differences in model structure show up on the hardware.
Model form and hardware fit
| CNN | Transformer / VLM | |
|---|---|---|
| Dominant operation | Convolution (local MACs) | Huge matrix products, softmax, normalization |
| Data reuse | High. Weight sharing pays off | Low. During generation especially, the weights are re-read every time |
| Bottleneck | Leans toward compute | Leans toward memory bandwidth |
| Memory use | Essentially fixed | Grows with input length (KV cache) |
| Support on existing NPUs | Almost certainly supported | Depends on generation and SDK. Check before committing |
The KV cache, a new kind of problem
When a Transformer generates text, every token it emits refers back to the Keys and Values of every token before it. Recomputing them each time would be wasteful, so they are kept around. This is the KV cache
.
Comments
Sign in to comment