Key facts
| What it stores | Key and value tensors per layer for every token in context |
| Size formula | Two tensors times layers times KV heads times head dimension times tokens times bytes |
| Growth | Linear in context length and batch size |
| Grouped-query attention | Fewer KV heads than attention heads, which shrinks the cache |
| Quantization | 8-bit or 4-bit KV cache cuts its memory roughly in half or to a quarter |
| Paged attention | Allocates cache in blocks to reduce fragmentation and raise throughput |
| Cloud option | Plugsky serves long-context models over an OpenAI-compatible API |
TL;DR
- The KV cache holds past token states so each new token costs one forward pass, not a full replay.
- Its memory grows linearly with context length and concurrency.
- Grouped-query attention is the main architectural lever that shrinks it.
- Quantizing the KV cache is often cheaper than shrinking weights further.
- Cap context and batch size to keep memory predictable.
How it works, step by step
- Estimate cache size with the formula for your model's layer and head configuration.
- Add weights, cache and runtime overhead to get true memory use at your context length.
- Check whether the model uses grouped-query or multi-query attention.
- Reduce context, batch size or concurrent sequences if memory is tight.
- Enable KV cache quantization in your runtime and re-measure quality.
- Test long-context tasks after any cache change to catch quality regressions.
- Plan for worst-case concurrency, not average, when sizing hardware.
Try it yourself
Open the KV cache calculator →
What the cache does and how it is sized
Transformers generate text one token at a time, and attention needs the key and value vectors of every earlier token. Recomputing them each step would be quadratic and slow, so runtimes store them. That store is the KV cache.
Its size is predictable: two tensors, key and value, for every layer of the model, for every token in the context. Each tensor has one entry per key-value head with head dimension values. Multiply by the number of tokens and the bytes per value, and you have the cache size. A model with 32 layers, 8 KV heads and head dimension 128 using 16-bit values needs about 128 KB per token, so 32k tokens consume roughly 4 GB before batch size multiplies it further.
Why long context and concurrency get expensive
Two multipliers catch teams out. The first is context length: doubling tokens doubles cache memory. The second is batch size: serving multiple requests at once keeps a separate cache per sequence, so concurrency multiplies memory again. That is why a model that runs comfortably in a single-user chat can run out of memory under a modest load.
- Long documents and code files inflate context even when the answer is short.
- Chat history accumulates unless you summarise or truncate.
- Long outputs extend the cache as generation proceeds.
- Retrieval prompts with many chunks consume cache proportional to total tokens.
Paged attention, used by high-throughput serving stacks, allocates cache in fixed-size blocks so memory is not wasted on fragmentation and can be shared across sequences with common prefixes.
Reducing cache cost
Architecture helps first. Grouped-query attention reduces the number of KV heads relative to attention heads, shrinking the cache by the ratio between them, often four to eight times. Multi-query attention goes further. When choosing a model family, this matters as much as parameter count for long-context work.
Then come runtime levers: KV cache quantization to 8-bit or 4-bit, smaller context windows, lower concurrency, and prefix caching to reuse shared prefixes across requests. Each trades some quality or flexibility for memory. Measure on your own long-context prompts, because cache quantization effects show up first on tasks that depend on precision.
If managing cache memory is not the point of your project, a hosted API absorbs the problem. Plugsky serves long-context and 30+ other models behind an OpenAI-compatible endpoint with chat, streaming, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. See the live pricing page for plans.
Honest comparison
| Lever | Memory effect | Quality effect | Implementation |
|---|---|---|---|
| Grouped-query attention | Large reduction versus multi-head | Small | Model architecture choice |
| KV cache quantization | Half at 8-bit, quarter at 4-bit | Workload-dependent | Runtime flag |
| Shorter context | Linear reduction | Less information available | Prompt and truncation policy |
| Lower batch size | Linear reduction | Lower throughput | Server configuration |
| Paged attention | Reduces fragmentation, not size | None | Serving stack feature |
Frequently asked questions
Is the KV cache the same as the context window?
No. The context window is the maximum number of tokens the model can attend to; the KV cache is the memory used to store their key and value states.
Why does my model fit at 4k context but fail at 32k?
Weights memory is fixed, but cache memory grows linearly with tokens. At 32k the cache can be several gigabytes, pushing total memory past the GPU limit.
What is grouped-query attention?
A design where multiple attention heads share fewer key-value heads, which shrinks the KV cache substantially with little quality loss. Most modern open models use it.
Can I quantize the KV cache independently of weights?
Yes. Many runtimes let you keep 16-bit weights and quantize the cache to 8-bit or 4-bit, which is useful for long-context deployments.
Does the cache survive between requests?
Usually not by default. Prefix caching can reuse a shared prefix across requests, which helps when many prompts start with the same system message.
How does batch size affect cache memory?
Each concurrent sequence needs its own cache, so memory scales with concurrency. That is why throughput-oriented servers cap the number of active sequences.
Do hosted APIs expose KV cache controls?
Generally not; the provider manages memory. Plugsky handles this internally while serving 30+ models over an OpenAI-compatible API.