Local AI

What is the KV cache and why does it use so much memory?

The KV cache stores the key and value tensors for every token the model has already processed, so it does not recompute them on each new token. Its size grows with context length, batch size, layer count and attention head dimensions. That is why long conversations and long documents consume memory far beyond the model weights themselves.

Key facts

What it storesKey and value tensors per layer for every token in context
Size formulaTwo tensors times layers times KV heads times head dimension times tokens times bytes
GrowthLinear in context length and batch size
Grouped-query attentionFewer KV heads than attention heads, which shrinks the cache
Quantization8-bit or 4-bit KV cache cuts its memory roughly in half or to a quarter
Paged attentionAllocates cache in blocks to reduce fragmentation and raise throughput
Cloud optionPlugsky serves long-context models over an OpenAI-compatible API

TL;DR

  • The KV cache holds past token states so each new token costs one forward pass, not a full replay.
  • Its memory grows linearly with context length and concurrency.
  • Grouped-query attention is the main architectural lever that shrinks it.
  • Quantizing the KV cache is often cheaper than shrinking weights further.
  • Cap context and batch size to keep memory predictable.

How it works, step by step

  1. Estimate cache size with the formula for your model's layer and head configuration.
  2. Add weights, cache and runtime overhead to get true memory use at your context length.
  3. Check whether the model uses grouped-query or multi-query attention.
  4. Reduce context, batch size or concurrent sequences if memory is tight.
  5. Enable KV cache quantization in your runtime and re-measure quality.
  6. Test long-context tasks after any cache change to catch quality regressions.
  7. Plan for worst-case concurrency, not average, when sizing hardware.
1Estimate cache sizewith the formulafor your model's2Add weights, cacheand runtimeoverhead to get3Check whether themodel usesgrouped-query or4Reduce context,batch size orconcurrent5Enable KV cachequantization inyour runtime and6Test long-contexttasks after anycache change to

Try it yourself

Open the KV cache calculator →

What the cache does and how it is sized

Transformers generate text one token at a time, and attention needs the key and value vectors of every earlier token. Recomputing them each step would be quadratic and slow, so runtimes store them. That store is the KV cache.

Its size is predictable: two tensors, key and value, for every layer of the model, for every token in the context. Each tensor has one entry per key-value head with head dimension values. Multiply by the number of tokens and the bytes per value, and you have the cache size. A model with 32 layers, 8 KV heads and head dimension 128 using 16-bit values needs about 128 KB per token, so 32k tokens consume roughly 4 GB before batch size multiplies it further.

Why long context and concurrency get expensive

Two multipliers catch teams out. The first is context length: doubling tokens doubles cache memory. The second is batch size: serving multiple requests at once keeps a separate cache per sequence, so concurrency multiplies memory again. That is why a model that runs comfortably in a single-user chat can run out of memory under a modest load.

  • Long documents and code files inflate context even when the answer is short.
  • Chat history accumulates unless you summarise or truncate.
  • Long outputs extend the cache as generation proceeds.
  • Retrieval prompts with many chunks consume cache proportional to total tokens.

Paged attention, used by high-throughput serving stacks, allocates cache in fixed-size blocks so memory is not wasted on fragmentation and can be shared across sequences with common prefixes.

Reducing cache cost

Architecture helps first. Grouped-query attention reduces the number of KV heads relative to attention heads, shrinking the cache by the ratio between them, often four to eight times. Multi-query attention goes further. When choosing a model family, this matters as much as parameter count for long-context work.

Then come runtime levers: KV cache quantization to 8-bit or 4-bit, smaller context windows, lower concurrency, and prefix caching to reuse shared prefixes across requests. Each trades some quality or flexibility for memory. Measure on your own long-context prompts, because cache quantization effects show up first on tasks that depend on precision.

If managing cache memory is not the point of your project, a hosted API absorbs the problem. Plugsky serves long-context and 30+ other models behind an OpenAI-compatible endpoint with chat, streaming, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. See the live pricing page for plans.

Honest comparison

LeverMemory effectQuality effectImplementation
Grouped-query attentionLarge reduction versus multi-headSmallModel architecture choice
KV cache quantizationHalf at 8-bit, quarter at 4-bitWorkload-dependentRuntime flag
Shorter contextLinear reductionLess information availablePrompt and truncation policy
Lower batch sizeLinear reductionLower throughputServer configuration
Paged attentionReduces fragmentation, not sizeNoneServing stack feature

Frequently asked questions

Is the KV cache the same as the context window?

No. The context window is the maximum number of tokens the model can attend to; the KV cache is the memory used to store their key and value states.

Why does my model fit at 4k context but fail at 32k?

Weights memory is fixed, but cache memory grows linearly with tokens. At 32k the cache can be several gigabytes, pushing total memory past the GPU limit.

What is grouped-query attention?

A design where multiple attention heads share fewer key-value heads, which shrinks the KV cache substantially with little quality loss. Most modern open models use it.

Can I quantize the KV cache independently of weights?

Yes. Many runtimes let you keep 16-bit weights and quantize the cache to 8-bit or 4-bit, which is useful for long-context deployments.

Does the cache survive between requests?

Usually not by default. Prefix caching can reuse a shared prefix across requests, which helps when many prompts start with the same system message.

How does batch size affect cache memory?

Each concurrent sequence needs its own cache, so memory scales with concurrency. That is why throughput-oriented servers cap the number of active sequences.

Do hosted APIs expose KV cache controls?

Generally not; the provider manages memory. Plugsky handles this internally while serving 30+ models over an OpenAI-compatible API.