Key facts
| VRAM | 8 GB GDDR6, shared with the display and other GPU apps |
| Comfortable range | 3B-8B models at 4-bit fit with modest context |
| Weights example | A 7B model at 4-bit is roughly 4-5 GB before KV cache |
| KV cache | Grows with context length, layer count and KV head count |
| Offload trade-off | Spilling layers to system RAM slows generation sharply |
| Runtime options | Ollama, LM Studio and llama.cpp support CUDA and partial offload |
| Cloud fallback | Plugsky exposes 30+ models through an OpenAI-compatible API |
| Endpoint status | Chat, streaming, tools, JSON mode and embeddings are live |
TL;DR
- 8 GB of VRAM caps you at small, 4-bit models for real work.
- Cap context length; the KV cache is what usually triggers out-of-memory errors.
- Keep the model fully on the GPU, because partial offload is a fallback, not a plan.
- Use CUDA-capable runtimes and watch dedicated GPU memory while testing.
- Send oversized jobs to a hosted OpenAI-compatible API instead of accepting CPU spills.
How it works, step by step
- Install current NVIDIA drivers and a CUDA-capable runtime such as Ollama, LM Studio or llama.cpp.
- Start with a 7B-8B model at 4-bit and a 4K-8K context window.
- Confirm all layers are on the GPU and check dedicated VRAM usage.
- Increase context only until memory is near the limit, then stop.
- Measure generation speed with your real prompt lengths, not short demos.
- Try a smaller model or higher quantization if quality suffers at 4-bit.
- Route heavy or long-context tasks to an OpenAI-compatible cloud endpoint.
Original data
Try it yourself
What 8 GB of VRAM actually buys you
Model weights are the first charge on VRAM. At 4-bit quantization, a 7B-8B model occupies roughly 4-5 GB, a 3B-4B model around 2-3 GB, and a 13B model about 7-8 GB before any context. That leaves little room for the KV cache on a 13B model, which is why 8 GB users generally settle on the 7B-8B class.
Windows and your browser also consume GPU memory. Before loading a model, check dedicated GPU memory in Task Manager; if the desktop is already using a gigabyte, your effective budget is smaller than the card's headline number.
Sizing models, quantization and context
Quantization trades a small amount of quality for large memory savings. Q4_K_M is the common default; Q5_K_M or Q6_K help when the model still fits; Q8_0 is near-lossless but roughly doubles the 4-bit footprint. Try both on your own evaluation set rather than trusting generic quality claims.
The KV cache is the variable that surprises people. It scales with context length, the number of layers and the number of key/value heads, so a model that runs at 4K context may fail at 32K. Practical levers:
- Lower the context window to what your prompts actually use.
- Reduce batch size to one for interactive use.
- Enable KV cache quantization if your runtime supports it.
- Close overlays, browsers and games that hold GPU memory.
Offload, thermals and when to use the cloud
llama.cpp can split a model between GPU and CPU, but every layer that lands in system memory slows token generation, because system RAM bandwidth is far lower than VRAM. If a model only just fits with offload, a smaller quantized model fully on the GPU usually feels better in daily use.
A 4060 class card is silent and efficient for single-user chat, coding assistance and light RAG. It is not built for concurrent users or very long documents. For those workloads, an OpenAI-compatible hosted endpoint keeps the same client code: Plugsky serves chat, streaming, tools, JSON mode, embeddings and RAG live, with audio, image, batch and fine-tuning endpoints coming soon. See pricing for plans and start free with plugsky-micro and plugsky-lite.
Honest comparison
| Concern | RTX 4060 8 GB | Plugsky hosted API | Check before deciding |
|---|---|---|---|
| Model size | 3B-8B at 4-bit fits well | 30+ models regardless of local memory | Largest model your task needs |
| Context length | Limited by KV cache at 8K-32K | Model-dependent context windows | Longest prompt you actually use |
| Concurrency | One user, one model at a time | Scales with the service | Concurrent requests at peak |
| Data path | Nothing leaves the machine | Requests go to your deployment | Data classification rules |
| Cost shape | Hardware, power and time | Flat monthly plans | Utilisation and budget |
Frequently asked questions
Can the RTX 4060 run a 13B model?
Only with aggressive quantization and partial offload. A 13B model at 4-bit is roughly 7-8 GB before KV cache, so it will not fit comfortably in 8 GB and layers will spill to system RAM.
Why do I get out-of-memory errors with a 7B model?
The weights fit, but the KV cache grows with context length and concurrency. Lower the context window, close other GPU apps or reduce batch size.
Is CUDA required?
No, llama.cpp can run on CPU, but CUDA is what makes the GPU fast. Install current NVIDIA drivers and use a runtime built with CUDA support.
How much context can 8 GB handle?
It depends on the model's layers and KV head configuration. Start at 4K, raise it in steps and watch VRAM; the KV cache, not the weights, is usually the limit.
Should I quantize the KV cache?
On runtimes that support it, KV cache quantization extends usable context at some quality cost. Test it against your eval before adopting it.
When should I stop optimising locally?
When the model you need will not fit, when you need concurrency, or when long prompts dominate. A hosted OpenAI-compatible endpoint covers those cases without a code rewrite.
Does the 4060 work for local RAG?
Yes for small corpora and modest context. Embeddings are cheap, but retrieval-augmented prompts are long, so budget the KV cache carefully.