Local AI

What local AI models can you run with 12GB of VRAM?

With 12GB of VRAM, comfortable choices are 7B-9B models at 4-bit to 6-bit precision, and 12B-14B models at 4-bit with room for modest context. Quantized 8B models at Q8 fit too, but leave little cache headroom. Larger models work with partial CPU offload at reduced speed. The practical rule is to leave at least a few gigabytes for the KV cache and runtime.

Key facts

Weights at 4-bit7B is roughly 4-5 GB, 14B roughly 8-9 GB
Weights at 8-bit7B-9B models land near 7-10 GB
Cache costLong context adds hundreds of megabytes to several gigabytes
Comfort zone7B-14B models at 4-bit with 4k-8k context
Offloadllama.cpp can split layers between GPU and CPU when weights exceed VRAM
Common GPUs12GB cards include the RTX 3060 12GB and 4070-class parts
Cloud optionPlugsky serves 30+ hosted models when local memory limits quality

TL;DR

  • 12GB is a strong 7B-9B tier and a workable 14B tier at 4-bit.
  • Reserve several gigabytes for KV cache; weights alone do not fit the budget.
  • Prefer Q4 or Q5 quantizations; Q8 fits small models but limits context.
  • Partial offload keeps bigger models usable at lower speed.
  • A 14B model at 4-bit usually beats a 7B model at 8-bit for the same memory.

How it works, step by step

  1. Work out weight memory: parameters times bits per weight divided by eight.
  2. Subtract that from 12GB and reserve 2-4 GB for KV cache and overhead.
  3. Pick a quantization level that fits the remainder with your target context.
  4. Serve with llama.cpp, Ollama or LM Studio and set a context cap.
  5. Measure tokens per second and memory use on a long prompt.
  6. If the model spills, reduce context, lower the quantization level or pick a smaller model.
  7. Route tasks that need a bigger model to a hosted API rather than fighting memory.
1Work out weightmemory: parameterstimes bits per2Subtract that from12GB and reserve2-4 GB for KV cache3Pick a quantizationlevel that fits theremainder with your4Serve withllama.cpp, Ollamaor LM Studio and5Measure tokens persecond and memoryuse on a long6If the modelspills, reducecontext, lower the

Original data

7B is roughly Weights at 4-bit7B-9B models lWeights at 8-bit7B-14B models Comfort zone12GB cards incCommon GPUsPlugsky servesCloud optionSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the LLM VRAM calculator →

What fits in 12GB

Weights are predictable: parameters times bits per weight. A 7B model at 4-bit is roughly 4-5 GB, at 8-bit roughly 7-8 GB, and at 16-bit around 14 GB, which does not fit. A 14B model at 4-bit lands near 8-9 GB, leaving a few gigabytes for cache. A 32B model at 4-bit needs roughly 17-19 GB, so it does not fit without offload.

That makes 7B-9B models the comfortable tier and 12B-14B models at 4-bit the practical upper tier. If you want a 14B model at higher precision, expect to exceed 12GB and plan for partial offload or a hosted fallback.

Context, cache and offload

The KV cache is what turns a fitting model into an out-of-memory error. Long prompts, long outputs and larger context windows all consume cache memory, and every concurrent request adds another sequence. Keep your runtime's context cap honest: choosing 32k when you only need 8k wastes memory that could hold a better model.

  • 4k-8k context: fine for chat, drafting and most RAG prompts.
  • 16k+ context: feasible for 7B-9B models at 4-bit with cache-fitting.
  • Offload: llama.cpp can move layers to CPU, trading speed for capacity.
  • Cache quantization: 8-bit cache roughly halves cache memory.

Offload is a useful escape hatch, not a strategy. A model with half its layers on CPU runs several times slower, which matters most for interactive use.

Choosing models and setting expectations

Pick task first. For coding assistance, a 7B-14B code-specialised model at 4-bit is a good fit. For chat and summarisation, a general instruction model at 7B-9B works well. For RAG, prioritise context handling over raw size, and use a local or hosted embedding model for retrieval. For agents, prefer reliable tool calling over parameter count, because loop iterations multiply any per-call weakness.

Set realistic expectations: 12GB covers private, everyday workloads well but caps model quality. When a task needs a frontier model, long context or high concurrency, route it to a hosted API. Plugsky serves 30+ models behind an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. The free plan includes plugsky-micro and plugsky-lite, and a 14-day full-access trial covers evaluation. Check the live pricing page for plans.

Honest comparison

Model size4-bit weights8-bit weightsVerdict for 12GB
3B-4BAbout 2-2.5 GBAbout 4-5 GBVery comfortable, long context possible
7B-9BAbout 4-6 GBAbout 7-10 GBComfortable at 4-bit, tight at 8-bit
12B-14BAbout 8-9 GBAbout 13-15 GBFits at 4-bit with modest context
30B-32BAbout 17-19 GBDoes not fitNeeds offload or more VRAM

Frequently asked questions

Can I run a 13B model on 12GB VRAM?

Yes, at 4-bit quantization with a modest context window. Reserve memory for the KV cache and cap context length to avoid out-of-memory errors.

Is 4-bit quality enough at this size?

For most drafting, summarisation and chat work, yes. Test your specific task; if quality falls short, try Q5 or Q6 before switching to a different model.

How much context can I use?

For a 7B-9B model at 4-bit, 8k-16k tokens is realistic. For 12B-14B models, stay closer to 4k-8k unless you quantize the KV cache.

Should I use CPU offload?

Only when you need the model and can accept slower generation. Offload turns a fast model into an interactive-latency problem.

What about MoE models?

Mixture-of-experts models keep all weights in memory but activate a subset per token, so they can be fast, yet memory still has to hold the full model. Check total size, not active size.

Can I run multiple models at once?

Small models can share a 12GB card, but switching costs load time and splitting memory usually reduces quality for the active task. Run one model well.

What if my workload needs a bigger model?

Use a hybrid route: keep 12GB local for routine work and send heavier tasks to a hosted OpenAI-compatible API such as Plugsky.