Key facts
| Weights at 4-bit | 3B is roughly 1.5-2 GB, 7B-8B roughly 4-5 GB |
| Weights at 8-bit | 3B-4B models land near 3-4 GB |
| Cache cost | Even short contexts need hundreds of megabytes |
| Comfort zone | 3B-4B at 4-8 bit, 7B-8B at 4-bit with 4k context |
| Offload | llama.cpp can move layers to CPU when weights exceed VRAM |
| Common GPUs | 8GB cards include RTX 4060-class and laptop GPUs |
| Cloud option | Plugsky hosts 30+ models when local memory is the bottleneck |
TL;DR
- 8GB is a small-model tier: 3B-4B models run well, 7B-8B models run tightly at 4-bit.
- Keep context short; the KV cache is what pushes a fitting model over the limit.
- Prefer Q4 quantizations and leave 1-2 GB of headroom.
- CPU offload makes bigger models possible but several times slower.
- Route demanding tasks to a hosted API instead of fighting for every megabyte.
How it works, step by step
- Compute weight memory at 4-bit for the models you want to try.
- Subtract from 8GB and reserve 1-2 GB for cache and overhead.
- Pick a quantization level that fits with your target context length.
- Serve with llama.cpp, Ollama or LM Studio and cap context explicitly.
- Test with a long prompt to confirm memory headroom, not just a short one.
- If the model spills, lower context or quantization before switching models.
- Use a hosted fallback for larger models and heavier reasoning.
Original data
Try it yourself
Open the LLM VRAM calculator →
The realistic model menu
At 4-bit, a 3B model needs roughly 1.5-2 GB, a 4B model around 2-2.5 GB and a 7B-8B model roughly 4-5 GB. With 8GB total, that leaves room for cache only at the smaller end. A 7B-8B model at 4-bit fits, but context must stay modest, typically 4k tokens or less depending on architecture and cache precision.
At 8-bit, 3B-4B models are the practical class; a 7B model at 8-bit needs roughly 7-8 GB and leaves almost nothing for cache. That is why 4-bit is the default at this tier, and why Q4_K_M-style levels are common. Higher precision only makes sense when the model is small enough to leave real headroom.
Context, cache and offload tactics
The KV cache is the usual culprit when a model that seemed to fit starts throwing memory errors. A 7B-8B model with grouped-query attention costs tens to hundreds of kilobytes per token depending on configuration, so a few thousand tokens can consume most of the headroom you reserved. Long chat histories and large retrieved passages have the same effect.
- Cap context in the runtime rather than trusting the model's maximum.
- Summarise or truncate chat history before it grows.
- Retrieve fewer, better chunks for RAG prompts.
- Quantize the KV cache to 8-bit if the runtime supports it.
- Use CPU offload only when you can accept slower generation.
Offload turns an interactive assistant into a batch tool. It is a legitimate choice for overnight jobs, not for chat.
Choosing the right small model
Task fit beats size at this tier. For chat and summarisation, a 3B-4B instruction model at 4-bit to 6-bit gives smooth interaction. For coding, a small code-specialised model at 4-bit handles completion and short explanations. For RAG, small models can answer well when retrieval is tight, but long contexts will not fit. For agents, tool-calling reliability matters more than parameter count, so test structured output carefully.
Accept the ceiling and plan around it. An 8GB machine is excellent for private, low-latency, occasional work, but it cannot host frontier quality or long context. Sending those requests to a hosted API keeps the workflow intact. Plugsky serves 30+ models behind an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. The free plan includes plugsky-micro and plugsky-lite, and a 14-day full-access trial covers evaluation. Check the live pricing page for plans.
Honest comparison
| Model size | 4-bit weights | 8-bit weights | Verdict for 8GB |
|---|---|---|---|
| 1.5B-2B | About 1-1.5 GB | About 2-3 GB | Very comfortable, long context possible |
| 3B-4B | About 1.5-2.5 GB | About 3-4 GB | Comfortable with moderate context |
| 7B-8B | About 4-5 GB | About 7-8 GB | Fits at 4-bit with short context |
| 13B+ | About 7-8 GB+ | Does not fit | Needs offload; slow generation |
Frequently asked questions
Can I run a 7B model on 8GB VRAM?
Yes, at 4-bit quantization with a short context window. Keep prompts and history lean, and avoid long retrieved passages.
What is the best model size for 8GB?
3B-4B models at 4-bit to 6-bit give the smoothest experience. They handle chat, summarisation and light coding well and leave room for context.
Why does my model run out of memory mid-conversation?
The KV cache grows with each token, so a long conversation can exceed the headroom you had at the start. Cap context or summarise history.
Does CPU offload help?
It lets a larger model run by moving layers to system memory, but generation slows substantially. Use it for batch work, not interactive chat.
Should I use Q4 or Q5?
Start with Q4_K_M-level quantization at this tier. Q5 or Q6 only if the model is small enough to leave cache headroom.
Can I run a local coding assistant on 8GB?
Yes, with a small code-specialised model at 4-bit. Expect competent completion and short explanations rather than complex refactoring.
Can I serve multiple users with 8GB?
Not comfortably. Cache memory multiplies per sequence, so concurrency exhausts VRAM quickly. Hosted inference is better for multi-user traffic.
What if I need a bigger model?
Keep the local model for routine work and route heavy tasks to a hosted OpenAI-compatible API such as Plugsky, which serves 30+ models.