Key facts
| Weights at 4-bit | 32B is roughly 17-19 GB, 14B roughly 8-9 GB |
| Weights at 8-bit | 14B is roughly 14-15 GB, 32B does not fit |
| Weights at 16-bit | 7B-9B models land near 14-18 GB |
| Cache cost | Long context can add several gigabytes; reserve 3-5 GB |
| Comfort zone | 32B at 4-bit with 8k-16k context |
| Multi-GPU | Two 24GB cards pool memory for larger models, with interconnect caveats |
| Cloud option | Plugsky serves 30+ hosted models for frontier and high-concurrency work |
TL;DR
- 24GB unlocks 30B-32B models at 4-bit, the biggest quality jump in the consumer tiers.
- It also runs 14B models at 8-bit, a strong quality-speed balance.
- 70B models need more memory than one 24GB card can provide.
- Reserve 3-5 GB for cache at longer contexts before choosing a quantization.
- Use hosted inference for frontier reasoning, long context and concurrent traffic.
How it works, step by step
- Compute weight memory for 32B and 14B classes at your candidate quantization levels.
- Reserve 3-5 GB for KV cache and runtime overhead.
- Choose between a 32B model at 4-bit and a 14B model at 8-bit based on task.
- Set a context cap that matches the workload, then benchmark at that length.
- For multi-GPU, verify runtime support and interconnect before buying a second card.
- Serve with vLLM for shared access or llama.cpp for single-user control.
- Route frontier-quality and long-context work to a hosted API when needed.
Original data
Try it yourself
The 24GB sweet spot
At 24GB, model choice opens up. A 32B model at 4-bit needs roughly 17-19 GB of weights, leaving a few gigabytes for cache and runtime. A 14B model at 8-bit needs roughly 14-15 GB and preserves more precision. A 7B-9B model at 16-bit fits with room for long context. This is the first tier where 30B-class quality becomes a daily option rather than an experiment.
The trade-off is context. Long prompts and concurrent requests consume cache, so a 32B model that fits at 8k tokens may fail at 32k. Quantize the KV cache to 8-bit if long context matters, or choose a 14B model at higher precision when context is the priority.
Serving, batching and multi-GPU
Single-user work runs well on llama.cpp, Ollama or LM Studio. For shared access, vLLM's batching makes better use of the card, but memory is divided between weights and cache per sequence, so cap concurrency deliberately. A 32B model leaves less cache room than a 14B model, which means fewer parallel requests before latency degrades.
- 32B at 4-bit: best local quality, modest context and concurrency.
- 14B at 8-bit: strong balance for interactive and multi-user use.
- 8B at 16-bit: fastest and most precise at small size, ideal for coding and tools.
Two 24GB cards can pool memory for 70B-class models at 4-bit, but interconnect bandwidth and software support determine whether the speed is acceptable. Measure before assuming dual-GPU doubles performance.
Knowing the ceiling and routing the rest
24GB covers private chat, serious RAG, coding help, and agents with real context. It does not cover 70B-class quality at speed, very long context at high concurrency, or frontier reasoning. Those tasks belong on a hosted API unless you buy server hardware and accept its operating costs.
Hybrid routing is the practical answer: keep daily and private workloads local and send the rest upstream. Plugsky serves 30+ models behind an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live, and supports cloud, VPC, on-prem and air-gapped deployment for enterprises. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. The free plan includes plugsky-micro and plugsky-lite, with a 14-day full-access trial for evaluation. See the live pricing page for tiers.
Honest comparison
| Model size | 4-bit weights | 8-bit weights | 16-bit weights | Verdict for 24GB |
|---|---|---|---|---|
| 7B-9B | About 4-6 GB | About 7-10 GB | About 14-18 GB | All fit; 16-bit with long context |
| 12B-14B | About 8-9 GB | About 14-15 GB | Does not fit | 4-bit or 8-bit both viable |
| 30B-32B | About 17-19 GB | Does not fit | Does not fit | Fits at 4-bit with modest context |
| 70B | About 35-40 GB | Does not fit | Does not fit | Needs two GPUs or aggressive quantization |
Frequently asked questions
Can I run a 70B model on 24GB?
Not without offload or very low-bit quantization that degrades quality. Two 24GB GPUs can pool memory for 70B models at 4-bit, but performance depends on interconnect and runtime support.
Should I choose a 32B model at 4-bit or a 14B at 8-bit?
It depends on the task. The 32B model generally reasons better; the 14B at 8-bit is faster, supports more context and suits interactive use.
How much context fits with a 32B model?
Roughly 8k-16k tokens with 4-bit weights, depending on the model's architecture and cache precision. Quantize the KV cache for longer context.
Is a 24GB GPU enough for multiple users?
For light concurrent use, yes, if you cap context and batch size. Heavy multi-user traffic needs batching software and more memory headroom.
What about fine-tuning?
Full fine-tuning needs far more memory than inference. LoRA and QLoRA adapters are more realistic on 24GB, but even then memory depends on model size and sequence length.
Do I need special software for multi-GPU?
Yes. Your runtime must support tensor or pipeline parallelism for your model, and results vary by interconnect. Test with your exact stack before investing.
When is a hosted API the better choice?
For frontier reasoning, 128k-plus context, high concurrency or workloads that spike. Plugsky provides 30+ models behind the same OpenAI-compatible interface.