Key facts
| Total memory | Weights plus KV cache plus runtime overhead |
| Weights | Parameters times bits per weight divided by eight |
| Cache formula | Two times layers times KV heads times head dimension times tokens times bytes |
| Example | 32 layers, 8 KV heads, head dimension 128, 16-bit: about 128 KB per token |
| Growth factors | Context length and concurrent sequences multiply cache memory |
| Reduction levers | Grouped-query attention, cache quantization, shorter context, lower batch |
| Cloud option | Plugsky serves long-context models without local memory planning |
TL;DR
- Weights are fixed; cache scales with tokens and concurrency.
- At 16-bit, a typical 8B model costs about 128 KB of cache per token.
- A 32k-token session can add roughly 4 GB before batching.
- Cache quantization and grouped-query attention are the biggest levers.
- Leave 10-20% headroom for runtime overhead and fragmentation.
How it works, step by step
- Find the model's layer count, KV head count and head dimension.
- Compute bytes per token for the cache at your chosen cache precision.
- Multiply by your target context length to get the cache budget.
- Multiply by concurrent sequences to get worst-case cache use.
- Add weights memory at your quantization level plus runtime overhead.
- Compare the total against available VRAM or unified memory and adjust.
- Validate with a real load test, since actual usage depends on the runtime.
Try it yourself
Open the LLM context window calculator →
The three lines of the memory budget
Weights are the simplest line: parameter count times bits per weight. A 7B model at 4-bit needs roughly 4-5 GB including quantization overhead, and an 8B model at 16-bit needs about 16 GB. This number does not change with context.
The KV cache is the variable line. Every token in context stores a key and a value tensor for every layer. With grouped-query attention, those tensors exist per KV head rather than per attention head. The formula is two times layers times KV heads times head dimension times tokens times bytes per value.
Runtime overhead is the third line: activations, temporary buffers, tokenizer state and allocator fragmentation. A 10-20% margin is a reasonable planning allowance, and some runtimes reserve more for graph capture or workspace.
Worked examples
Take a model with 32 layers, 8 KV heads and head dimension 128. At 16-bit cache precision each token costs 2 x 32 x 8 x 128 x 2 bytes, which is about 128 KB. At 8k context that is roughly 1 GB; at 32k it is about 4 GB; at 128k it approaches 16 GB. Cache quantization to 8-bit halves each figure.
Now multiply by concurrency. Four simultaneous long requests at 32k context could require around 16 GB of cache alone. That is the scenario where a model that runs fine for one user fails under light load.
- Single-user chat: 4k-8k context, one sequence, cache is modest.
- Document QA: 16k-32k context, one or few sequences, plan 2-4 GB of cache.
- Batch summarisation: many sequences, cache dominates; cap concurrency explicitly.
Reducing the budget instead of buying hardware
You rarely need the theoretical maximum context. Summarise or trim chat history, retrieve fewer but better chunks, and cap output length. Each of these reduces tokens and cache directly. On the runtime side, enable KV cache quantization to 8-bit, lower the maximum concurrent sequences, and use prefix caching when requests share a system prompt.
Then verify with a load test at worst-case concurrency. Calculators give you the arithmetic; only a test shows the runtime's real overhead and how memory behaves as sequences grow. If long context is essential and local memory is not enough, a hosted API removes the constraint. Plugsky serves long-context models alongside 30+ others over an OpenAI-compatible endpoint, with chat, streaming, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. See the live pricing page for plan details.
Honest comparison
| Context length | Cache at 16-bit (example model) | Cache at 8-bit | Typical fit |
|---|---|---|---|
| 4k tokens | About 0.5 GB | About 0.25 GB | 8 GB GPUs |
| 8k tokens | About 1 GB | About 0.5 GB | 8-12 GB GPUs |
| 32k tokens | About 4 GB | About 2 GB | 16-24 GB GPUs |
| 128k tokens | About 16 GB | About 8 GB | 24 GB+ or unified memory |
Frequently asked questions
How much memory does context add to a model?
For a typical 8B-class model with grouped-query attention, roughly 128 KB per token at 16-bit cache precision, so 32k tokens adds about 4 GB.
Does prompt length affect memory before generation starts?
Yes. The prompt is processed first and populates the cache, so a long prompt consumes memory immediately.
What is the difference between context window and memory?
The context window is the maximum tokens the model can attend to. Memory is the physical storage required to hold those token states.
Why does memory grow during generation?
Each generated token is added to the cache, so memory increases as the output gets longer.
Can I exceed my VRAM and still run?
Some runtimes offload cache or layers to system memory, which prevents failure but slows generation significantly.
Does batching always multiply cache memory?
Each active sequence needs its own cache, so yes, concurrency scales cache memory. Paged attention reduces waste but not the fundamental size.
Where can I get the model's layer and head numbers?
Model configuration files, usually published with the weights, list layer count, attention heads and KV heads.
Do hosted APIs have the same limits?
Providers manage memory internally, so you usually just choose a model with the context length you need. Plugsky serves long-context models among its 30+ models over an OpenAI-compatible API.