Local AI

What local AI models can you run with 16GB of VRAM?

16GB of VRAM is the comfortable middle tier. It runs 7B-9B models at 8-bit, 12B-14B models at 4-bit to 6-bit, and small mixture-of-experts models that activate only part of their weights per token. Context of 8k-16k is realistic, and 32B-class models remain out of reach without offload. Reserve two to four gigabytes for the KV cache.

Key facts

Weights at 4-bit14B is roughly 8-9 GB, 24B roughly 13-14 GB
Weights at 8-bit7B-9B models land near 7-10 GB
Cache costKV cache grows with context and concurrency; reserve 2-4 GB
Comfort zone12B-14B at 4-6 bit with 8k-16k context
MoE modelsAll weights stay in memory while a subset activates per token
Common GPUs16GB cards include RTX 4060 Ti 16GB and workstation-class parts
Cloud optionPlugsky serves 30+ hosted models for tasks beyond local memory

TL;DR

  • 16GB covers 14B-class models at 4-bit and 8B models at higher precision.
  • It is the first tier where long context stops being painful.
  • Small MoE models give better speed per unit of quality but still need full weights in memory.
  • 32B models need offload, lower-bit quantization or more VRAM.
  • Keep a hosted fallback for frontier-quality tasks.

How it works, step by step

  1. Compute weight memory for the models you are considering at your quantization level.
  2. Reserve 2-4 GB for KV cache and runtime overhead before committing.
  3. Choose a model class: 8B at 8-bit for speed, 14B at 4-bit for quality.
  4. Set a context cap that matches your task, not the model's maximum.
  5. Benchmark tokens per second and memory at your real context length.
  6. Try a small MoE model if you want faster generation without shrinking quality.
  7. Route frontier tasks and spikes to a hosted OpenAI-compatible API.
1Compute weightmemory for themodels you are2Reserve 2-4 GB forKV cache andruntime overhead3Choose a modelclass: 8B at 8-bitfor speed, 14B at4Set a context capthat matches yourtask, not the5Benchmark tokensper second andmemory at your real6Try a small MoEmodel if you wantfaster generation

Original data

14B is roughlyWeights at 4-bit7B-9B models lWeights at 8-bitKV cache growsCache cost12B-14B at 4-6Comfort zone16GB cards incCommon GPUsPlugsky servesCloud optionSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the LLM VRAM calculator →

What fits comfortably

16GB changes the calculus from 12GB in two ways. First, 14B models at 4-bit fit with room for real context rather than a token budget. Second, 7B-9B models at 8-bit become practical, which preserves more quality than aggressive quantization. A 14B model at 4-bit needs roughly 8-9 GB of weights; an 8B model at 8-bit needs roughly 8 GB. Both leave several gigabytes for cache.

At 6-bit, a 14B model needs roughly 11-12 GB, which still fits but narrows the context budget. Above that, 24B and 32B models exceed 16GB at 4-bit, so they need offload, lower-bit quantization or a different machine.

Context, concurrency and MoE

With weights settled, cache dominates the remainder. A typical 8B-class model with grouped-query attention costs roughly 128 KB of KV cache per token at 16-bit precision, so 16k tokens consume about 2 GB. Cache quantization halves that, and concurrency multiplies it. If two users share the endpoint at 16k context, plan for roughly 4 GB of cache before overhead.

Mixture-of-experts models are worth attention at this tier. They keep all weights in memory but activate a fraction per token, which improves speed without reducing quality as much as quantization. The catch is memory: total parameters determine fit, active parameters determine compute. A small MoE can feel faster than a dense model of similar quality, at the cost of loading more weights.

Serving and when to go hybrid

For single-user workflows, llama.cpp, Ollama or LM Studio on a 16GB card is straightforward. For shared use, cap concurrency so cache stays within budget, and monitor memory rather than assuming weights are the only cost. If you serve multiple users or long contexts, consider vLLM for batching, but size the GPU for worst-case concurrency, not average.

Know the ceiling. 16GB is excellent for private chat, RAG, coding assistance and light agents, but frontier reasoning and very long context remain hosted work. A hybrid split keeps routine traffic local and sends hard tasks to an API. Plugsky serves 30+ models behind an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. Compare plans on the live pricing page.

Honest comparison

Model size4-bit weights6-bit weights8-bit weightsVerdict for 16GB
7B-9BAbout 4-6 GBAbout 6-8 GBAbout 7-10 GBAll comfortable, long context possible
12B-14BAbout 8-9 GBAbout 11-12 GBAbout 13-15 GB4-6 bit fits, 8-bit is tight
24BAbout 13-14 GBDoes not fitDoes not fitOnly at 4-bit with small context
32B+About 17-19 GBDoes not fitDoes not fitNeeds offload or more VRAM

Frequently asked questions

Can I run a 14B model on 16GB?

Yes, comfortably at 4-bit and workably at 6-bit. Reserve 2-4 GB for the KV cache and cap context at 8k-16k depending on the model.

Is 8-bit better than 4-bit at this tier?

8-bit preserves more quality but requires a smaller model to fit. Compare a 14B at 4-bit against an 8B at 8-bit on your own tasks; the larger model often wins.

How much context can I use?

For 8B-9B models, 16k tokens is realistic. For 14B models at 4-bit, stay near 8k unless you quantize the KV cache.

Are 32B models possible?

Not fully in 16GB at 4-bit. Partial offload can make them run, but speed drops substantially, which hurts interactive use.

What about mixture-of-experts models?

They can be a good fit because only a subset of weights activates per token, improving tokens per second. Memory must still hold the full model, so check total size.

Should I use vLLM on a 16GB card?

It works, but batch size is limited by memory. Cap concurrent sequences and context length so the cache does not exhaust VRAM.

What workloads suit 16GB best?

Private chat, document RAG, coding assistance and light tool-using agents. Heavy reasoning and long-context document analysis are better served by a hosted model.