Key facts
| Weights at 4-bit | 7B is roughly 4-5 GB, 14B roughly 8-9 GB |
| Weights at 8-bit | 7B-9B models land near 7-10 GB |
| Cache cost | Long context adds hundreds of megabytes to several gigabytes |
| Comfort zone | 7B-14B models at 4-bit with 4k-8k context |
| Offload | llama.cpp can split layers between GPU and CPU when weights exceed VRAM |
| Common GPUs | 12GB cards include the RTX 3060 12GB and 4070-class parts |
| Cloud option | Plugsky serves 30+ hosted models when local memory limits quality |
TL;DR
- 12GB is a strong 7B-9B tier and a workable 14B tier at 4-bit.
- Reserve several gigabytes for KV cache; weights alone do not fit the budget.
- Prefer Q4 or Q5 quantizations; Q8 fits small models but limits context.
- Partial offload keeps bigger models usable at lower speed.
- A 14B model at 4-bit usually beats a 7B model at 8-bit for the same memory.
How it works, step by step
- Work out weight memory: parameters times bits per weight divided by eight.
- Subtract that from 12GB and reserve 2-4 GB for KV cache and overhead.
- Pick a quantization level that fits the remainder with your target context.
- Serve with llama.cpp, Ollama or LM Studio and set a context cap.
- Measure tokens per second and memory use on a long prompt.
- If the model spills, reduce context, lower the quantization level or pick a smaller model.
- Route tasks that need a bigger model to a hosted API rather than fighting memory.
Original data
Try it yourself
Open the LLM VRAM calculator →
What fits in 12GB
Weights are predictable: parameters times bits per weight. A 7B model at 4-bit is roughly 4-5 GB, at 8-bit roughly 7-8 GB, and at 16-bit around 14 GB, which does not fit. A 14B model at 4-bit lands near 8-9 GB, leaving a few gigabytes for cache. A 32B model at 4-bit needs roughly 17-19 GB, so it does not fit without offload.
That makes 7B-9B models the comfortable tier and 12B-14B models at 4-bit the practical upper tier. If you want a 14B model at higher precision, expect to exceed 12GB and plan for partial offload or a hosted fallback.
Context, cache and offload
The KV cache is what turns a fitting model into an out-of-memory error. Long prompts, long outputs and larger context windows all consume cache memory, and every concurrent request adds another sequence. Keep your runtime's context cap honest: choosing 32k when you only need 8k wastes memory that could hold a better model.
- 4k-8k context: fine for chat, drafting and most RAG prompts.
- 16k+ context: feasible for 7B-9B models at 4-bit with cache-fitting.
- Offload: llama.cpp can move layers to CPU, trading speed for capacity.
- Cache quantization: 8-bit cache roughly halves cache memory.
Offload is a useful escape hatch, not a strategy. A model with half its layers on CPU runs several times slower, which matters most for interactive use.
Choosing models and setting expectations
Pick task first. For coding assistance, a 7B-14B code-specialised model at 4-bit is a good fit. For chat and summarisation, a general instruction model at 7B-9B works well. For RAG, prioritise context handling over raw size, and use a local or hosted embedding model for retrieval. For agents, prefer reliable tool calling over parameter count, because loop iterations multiply any per-call weakness.
Set realistic expectations: 12GB covers private, everyday workloads well but caps model quality. When a task needs a frontier model, long context or high concurrency, route it to a hosted API. Plugsky serves 30+ models behind an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. The free plan includes plugsky-micro and plugsky-lite, and a 14-day full-access trial covers evaluation. Check the live pricing page for plans.
Honest comparison
| Model size | 4-bit weights | 8-bit weights | Verdict for 12GB |
|---|---|---|---|
| 3B-4B | About 2-2.5 GB | About 4-5 GB | Very comfortable, long context possible |
| 7B-9B | About 4-6 GB | About 7-10 GB | Comfortable at 4-bit, tight at 8-bit |
| 12B-14B | About 8-9 GB | About 13-15 GB | Fits at 4-bit with modest context |
| 30B-32B | About 17-19 GB | Does not fit | Needs offload or more VRAM |
Frequently asked questions
Can I run a 13B model on 12GB VRAM?
Yes, at 4-bit quantization with a modest context window. Reserve memory for the KV cache and cap context length to avoid out-of-memory errors.
Is 4-bit quality enough at this size?
For most drafting, summarisation and chat work, yes. Test your specific task; if quality falls short, try Q5 or Q6 before switching to a different model.
How much context can I use?
For a 7B-9B model at 4-bit, 8k-16k tokens is realistic. For 12B-14B models, stay closer to 4k-8k unless you quantize the KV cache.
Should I use CPU offload?
Only when you need the model and can accept slower generation. Offload turns a fast model into an interactive-latency problem.
What about MoE models?
Mixture-of-experts models keep all weights in memory but activate a subset per token, so they can be fast, yet memory still has to hold the full model. Check total size, not active size.
Can I run multiple models at once?
Small models can share a 12GB card, but switching costs load time and splitting memory usually reduces quality for the active task. Run one model well.
What if my workload needs a bigger model?
Use a hybrid route: keep 12GB local for routine work and send heavier tasks to a hosted OpenAI-compatible API such as Plugsky.