Local AI

What local AI models can you run with 24GB of VRAM?

24GB of VRAM is the enthusiast and small-server tier. It runs 30B-32B models at 4-bit with usable context, 12B-14B models at 8-bit, and 7B-9B models at 16-bit. Larger models such as 70B remain out of reach without aggressive quantization or offload, and multi-GPU setups change that math. Reserve several gigabytes for the KV cache at longer contexts.

Key facts

Weights at 4-bit32B is roughly 17-19 GB, 14B roughly 8-9 GB
Weights at 8-bit14B is roughly 14-15 GB, 32B does not fit
Weights at 16-bit7B-9B models land near 14-18 GB
Cache costLong context can add several gigabytes; reserve 3-5 GB
Comfort zone32B at 4-bit with 8k-16k context
Multi-GPUTwo 24GB cards pool memory for larger models, with interconnect caveats
Cloud optionPlugsky serves 30+ hosted models for frontier and high-concurrency work

TL;DR

  • 24GB unlocks 30B-32B models at 4-bit, the biggest quality jump in the consumer tiers.
  • It also runs 14B models at 8-bit, a strong quality-speed balance.
  • 70B models need more memory than one 24GB card can provide.
  • Reserve 3-5 GB for cache at longer contexts before choosing a quantization.
  • Use hosted inference for frontier reasoning, long context and concurrent traffic.

How it works, step by step

  1. Compute weight memory for 32B and 14B classes at your candidate quantization levels.
  2. Reserve 3-5 GB for KV cache and runtime overhead.
  3. Choose between a 32B model at 4-bit and a 14B model at 8-bit based on task.
  4. Set a context cap that matches the workload, then benchmark at that length.
  5. For multi-GPU, verify runtime support and interconnect before buying a second card.
  6. Serve with vLLM for shared access or llama.cpp for single-user control.
  7. Route frontier-quality and long-context work to a hosted API when needed.
1Compute weightmemory for 32B and14B classes at your2Reserve 3-5 GB forKV cache andruntime overhead.3Choose between a32B model at 4-bitand a 14B model at4Set a context capthat matches theworkload, then5For multi-GPU,verify runtimesupport and6Serve with vLLM forshared access orllama.cpp for

Original data

32B is roughlyWeights at 4-bit14B is roughlyWeights at 8-bit7B-9B models lWeights at 16-bitLong context cCache cost32B at 4-bit wComfort zoneTwo 24GB cardsMulti-GPUSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the VRAM calculator →

The 24GB sweet spot

At 24GB, model choice opens up. A 32B model at 4-bit needs roughly 17-19 GB of weights, leaving a few gigabytes for cache and runtime. A 14B model at 8-bit needs roughly 14-15 GB and preserves more precision. A 7B-9B model at 16-bit fits with room for long context. This is the first tier where 30B-class quality becomes a daily option rather than an experiment.

The trade-off is context. Long prompts and concurrent requests consume cache, so a 32B model that fits at 8k tokens may fail at 32k. Quantize the KV cache to 8-bit if long context matters, or choose a 14B model at higher precision when context is the priority.

Serving, batching and multi-GPU

Single-user work runs well on llama.cpp, Ollama or LM Studio. For shared access, vLLM's batching makes better use of the card, but memory is divided between weights and cache per sequence, so cap concurrency deliberately. A 32B model leaves less cache room than a 14B model, which means fewer parallel requests before latency degrades.

  • 32B at 4-bit: best local quality, modest context and concurrency.
  • 14B at 8-bit: strong balance for interactive and multi-user use.
  • 8B at 16-bit: fastest and most precise at small size, ideal for coding and tools.

Two 24GB cards can pool memory for 70B-class models at 4-bit, but interconnect bandwidth and software support determine whether the speed is acceptable. Measure before assuming dual-GPU doubles performance.

Knowing the ceiling and routing the rest

24GB covers private chat, serious RAG, coding help, and agents with real context. It does not cover 70B-class quality at speed, very long context at high concurrency, or frontier reasoning. Those tasks belong on a hosted API unless you buy server hardware and accept its operating costs.

Hybrid routing is the practical answer: keep daily and private workloads local and send the rest upstream. Plugsky serves 30+ models behind an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live, and supports cloud, VPC, on-prem and air-gapped deployment for enterprises. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. The free plan includes plugsky-micro and plugsky-lite, with a 14-day full-access trial for evaluation. See the live pricing page for tiers.

Honest comparison

Model size4-bit weights8-bit weights16-bit weightsVerdict for 24GB
7B-9BAbout 4-6 GBAbout 7-10 GBAbout 14-18 GBAll fit; 16-bit with long context
12B-14BAbout 8-9 GBAbout 14-15 GBDoes not fit4-bit or 8-bit both viable
30B-32BAbout 17-19 GBDoes not fitDoes not fitFits at 4-bit with modest context
70BAbout 35-40 GBDoes not fitDoes not fitNeeds two GPUs or aggressive quantization

Frequently asked questions

Can I run a 70B model on 24GB?

Not without offload or very low-bit quantization that degrades quality. Two 24GB GPUs can pool memory for 70B models at 4-bit, but performance depends on interconnect and runtime support.

Should I choose a 32B model at 4-bit or a 14B at 8-bit?

It depends on the task. The 32B model generally reasons better; the 14B at 8-bit is faster, supports more context and suits interactive use.

How much context fits with a 32B model?

Roughly 8k-16k tokens with 4-bit weights, depending on the model's architecture and cache precision. Quantize the KV cache for longer context.

Is a 24GB GPU enough for multiple users?

For light concurrent use, yes, if you cap context and batch size. Heavy multi-user traffic needs batching software and more memory headroom.

What about fine-tuning?

Full fine-tuning needs far more memory than inference. LoRA and QLoRA adapters are more realistic on 24GB, but even then memory depends on model size and sequence length.

Do I need special software for multi-GPU?

Yes. Your runtime must support tensor or pipeline parallelism for your model, and results vary by interconnect. Test with your exact stack before investing.

When is a hosted API the better choice?

For frontier reasoning, 128k-plus context, high concurrency or workloads that spike. Plugsky provides 30+ models behind the same OpenAI-compatible interface.