JA EN

#context

1 articles

01 ·Inference & Serving·★ MEMBER·10 min read Prompt Caching and Context Design — One Prefix Rule That Moves Your Bill by an Order of Magnitude Are you paying to have the same system prompt re-read on every single request? Prompt caching only works on exact prefixes — and that one rule decides what goes where in your context. Why a single timestamp at the top wipes out everything below it, and how misreading the TTL can make caching 25% more expensive than not caching at all.