KV cache
i
KV cache
Stored information about earlier tokens. Reusing it makes each new generation step faster, but requires additional memory.
Grow a sequence and compare repeated attention work with and without cached keys and values.
Reuse previous K/V states
Attention work24
Without cache300
Work avoided92%
The cache trades memory for speed: earlier key/value projections are stored instead of recomputed for every new token.
12345678910111213141516171819202122232425262728293031323334353637383940414243444546474849505152535455565758596061626364
Current token 24
readsCached K/V slots 23
then writesOne new slot +1