Local AI

What local AI models can you run with 8GB of VRAM?

With 8GB of VRAM, 3B-4B models run comfortably at 4-bit to 8-bit, and 7B-8B models fit at 4-bit with a modest context window. Larger models need partial CPU offload, which slows generation. The key discipline is budgeting: leave one to two gigabytes for the KV cache and runtime, and keep context short enough that weights and cache both fit.

Key facts

Weights at 4-bit3B is roughly 1.5-2 GB, 7B-8B roughly 4-5 GB
Weights at 8-bit3B-4B models land near 3-4 GB
Cache costEven short contexts need hundreds of megabytes
Comfort zone3B-4B at 4-8 bit, 7B-8B at 4-bit with 4k context
Offloadllama.cpp can move layers to CPU when weights exceed VRAM
Common GPUs8GB cards include RTX 4060-class and laptop GPUs
Cloud optionPlugsky hosts 30+ models when local memory is the bottleneck

TL;DR

  • 8GB is a small-model tier: 3B-4B models run well, 7B-8B models run tightly at 4-bit.
  • Keep context short; the KV cache is what pushes a fitting model over the limit.
  • Prefer Q4 quantizations and leave 1-2 GB of headroom.
  • CPU offload makes bigger models possible but several times slower.
  • Route demanding tasks to a hosted API instead of fighting for every megabyte.

How it works, step by step

  1. Compute weight memory at 4-bit for the models you want to try.
  2. Subtract from 8GB and reserve 1-2 GB for cache and overhead.
  3. Pick a quantization level that fits with your target context length.
  4. Serve with llama.cpp, Ollama or LM Studio and cap context explicitly.
  5. Test with a long prompt to confirm memory headroom, not just a short one.
  6. If the model spills, lower context or quantization before switching models.
  7. Use a hosted fallback for larger models and heavier reasoning.
1Compute weightmemory at 4-bit forthe models you want2Subtract from 8GBand reserve 1-2 GBfor cache and3Pick a quantizationlevel that fitswith your target4Serve withllama.cpp, Ollamaor LM Studio and5Test with a longprompt to confirmmemory headroom,6If the modelspills, lowercontext or

Original data

3B is roughly Weights at 4-bit3B-4B models lWeights at 8-bit3B-4B at 4-8 bComfort zone8GB cards inclCommon GPUsPlugsky hosts Cloud optionSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the LLM VRAM calculator →

The realistic model menu

At 4-bit, a 3B model needs roughly 1.5-2 GB, a 4B model around 2-2.5 GB and a 7B-8B model roughly 4-5 GB. With 8GB total, that leaves room for cache only at the smaller end. A 7B-8B model at 4-bit fits, but context must stay modest, typically 4k tokens or less depending on architecture and cache precision.

At 8-bit, 3B-4B models are the practical class; a 7B model at 8-bit needs roughly 7-8 GB and leaves almost nothing for cache. That is why 4-bit is the default at this tier, and why Q4_K_M-style levels are common. Higher precision only makes sense when the model is small enough to leave real headroom.

Context, cache and offload tactics

The KV cache is the usual culprit when a model that seemed to fit starts throwing memory errors. A 7B-8B model with grouped-query attention costs tens to hundreds of kilobytes per token depending on configuration, so a few thousand tokens can consume most of the headroom you reserved. Long chat histories and large retrieved passages have the same effect.

  • Cap context in the runtime rather than trusting the model's maximum.
  • Summarise or truncate chat history before it grows.
  • Retrieve fewer, better chunks for RAG prompts.
  • Quantize the KV cache to 8-bit if the runtime supports it.
  • Use CPU offload only when you can accept slower generation.

Offload turns an interactive assistant into a batch tool. It is a legitimate choice for overnight jobs, not for chat.

Choosing the right small model

Task fit beats size at this tier. For chat and summarisation, a 3B-4B instruction model at 4-bit to 6-bit gives smooth interaction. For coding, a small code-specialised model at 4-bit handles completion and short explanations. For RAG, small models can answer well when retrieval is tight, but long contexts will not fit. For agents, tool-calling reliability matters more than parameter count, so test structured output carefully.

Accept the ceiling and plan around it. An 8GB machine is excellent for private, low-latency, occasional work, but it cannot host frontier quality or long context. Sending those requests to a hosted API keeps the workflow intact. Plugsky serves 30+ models behind an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. The free plan includes plugsky-micro and plugsky-lite, and a 14-day full-access trial covers evaluation. Check the live pricing page for plans.

Honest comparison

Model size4-bit weights8-bit weightsVerdict for 8GB
1.5B-2BAbout 1-1.5 GBAbout 2-3 GBVery comfortable, long context possible
3B-4BAbout 1.5-2.5 GBAbout 3-4 GBComfortable with moderate context
7B-8BAbout 4-5 GBAbout 7-8 GBFits at 4-bit with short context
13B+About 7-8 GB+Does not fitNeeds offload; slow generation

Frequently asked questions

Can I run a 7B model on 8GB VRAM?

Yes, at 4-bit quantization with a short context window. Keep prompts and history lean, and avoid long retrieved passages.

What is the best model size for 8GB?

3B-4B models at 4-bit to 6-bit give the smoothest experience. They handle chat, summarisation and light coding well and leave room for context.

Why does my model run out of memory mid-conversation?

The KV cache grows with each token, so a long conversation can exceed the headroom you had at the start. Cap context or summarise history.

Does CPU offload help?

It lets a larger model run by moving layers to system memory, but generation slows substantially. Use it for batch work, not interactive chat.

Should I use Q4 or Q5?

Start with Q4_K_M-level quantization at this tier. Q5 or Q6 only if the model is small enough to leave cache headroom.

Can I run a local coding assistant on 8GB?

Yes, with a small code-specialised model at 4-bit. Expect competent completion and short explanations rather than complex refactoring.

Can I serve multiple users with 8GB?

Not comfortably. Cache memory multiplies per sequence, so concurrency exhausts VRAM quickly. Hosted inference is better for multi-user traffic.

What if I need a bigger model?

Keep the local model for routine work and route heavy tasks to a hosted OpenAI-compatible API such as Plugsky, which serves 30+ models.