Local AI

How do you run local AI on an RTX 4060?

The RTX 4060 has 8 GB of VRAM, so run 3B-8B models at 4-bit quantization with a capped context window. Keep layers fully on the GPU, close other GPU-using apps, and watch the KV cache as context grows. When a task needs a bigger model, route it to an OpenAI-compatible cloud endpoint instead of offloading to slow system memory.

Key facts

VRAM8 GB GDDR6, shared with the display and other GPU apps
Comfortable range3B-8B models at 4-bit fit with modest context
Weights exampleA 7B model at 4-bit is roughly 4-5 GB before KV cache
KV cacheGrows with context length, layer count and KV head count
Offload trade-offSpilling layers to system RAM slows generation sharply
Runtime optionsOllama, LM Studio and llama.cpp support CUDA and partial offload
Cloud fallbackPlugsky exposes 30+ models through an OpenAI-compatible API
Endpoint statusChat, streaming, tools, JSON mode and embeddings are live

TL;DR

  • 8 GB of VRAM caps you at small, 4-bit models for real work.
  • Cap context length; the KV cache is what usually triggers out-of-memory errors.
  • Keep the model fully on the GPU, because partial offload is a fallback, not a plan.
  • Use CUDA-capable runtimes and watch dedicated GPU memory while testing.
  • Send oversized jobs to a hosted OpenAI-compatible API instead of accepting CPU spills.

How it works, step by step

  1. Install current NVIDIA drivers and a CUDA-capable runtime such as Ollama, LM Studio or llama.cpp.
  2. Start with a 7B-8B model at 4-bit and a 4K-8K context window.
  3. Confirm all layers are on the GPU and check dedicated VRAM usage.
  4. Increase context only until memory is near the limit, then stop.
  5. Measure generation speed with your real prompt lengths, not short demos.
  6. Try a smaller model or higher quantization if quality suffers at 4-bit.
  7. Route heavy or long-context tasks to an OpenAI-compatible cloud endpoint.
1Install currentNVIDIA drivers anda CUDA-capable2Start with a 7B-8Bmodel at 4-bit anda 4K-8K context3Confirm all layersare on the GPU andcheck dedicated4Increase contextonly until memoryis near the limit,5Measure generationspeed with yourreal prompt6Try a smaller modelor higherquantization if

Original data

8 GB GDDR6, shVRAM3B-8B models aComfortable rangeA 7B model at Weights examplePlugsky exposeCloud fallbackSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the VRAM calculator →

What 8 GB of VRAM actually buys you

Model weights are the first charge on VRAM. At 4-bit quantization, a 7B-8B model occupies roughly 4-5 GB, a 3B-4B model around 2-3 GB, and a 13B model about 7-8 GB before any context. That leaves little room for the KV cache on a 13B model, which is why 8 GB users generally settle on the 7B-8B class.

Windows and your browser also consume GPU memory. Before loading a model, check dedicated GPU memory in Task Manager; if the desktop is already using a gigabyte, your effective budget is smaller than the card's headline number.

Sizing models, quantization and context

Quantization trades a small amount of quality for large memory savings. Q4_K_M is the common default; Q5_K_M or Q6_K help when the model still fits; Q8_0 is near-lossless but roughly doubles the 4-bit footprint. Try both on your own evaluation set rather than trusting generic quality claims.

The KV cache is the variable that surprises people. It scales with context length, the number of layers and the number of key/value heads, so a model that runs at 4K context may fail at 32K. Practical levers:

  • Lower the context window to what your prompts actually use.
  • Reduce batch size to one for interactive use.
  • Enable KV cache quantization if your runtime supports it.
  • Close overlays, browsers and games that hold GPU memory.

Offload, thermals and when to use the cloud

llama.cpp can split a model between GPU and CPU, but every layer that lands in system memory slows token generation, because system RAM bandwidth is far lower than VRAM. If a model only just fits with offload, a smaller quantized model fully on the GPU usually feels better in daily use.

A 4060 class card is silent and efficient for single-user chat, coding assistance and light RAG. It is not built for concurrent users or very long documents. For those workloads, an OpenAI-compatible hosted endpoint keeps the same client code: Plugsky serves chat, streaming, tools, JSON mode, embeddings and RAG live, with audio, image, batch and fine-tuning endpoints coming soon. See pricing for plans and start free with plugsky-micro and plugsky-lite.

Honest comparison

ConcernRTX 4060 8 GBPlugsky hosted APICheck before deciding
Model size3B-8B at 4-bit fits well30+ models regardless of local memoryLargest model your task needs
Context lengthLimited by KV cache at 8K-32KModel-dependent context windowsLongest prompt you actually use
ConcurrencyOne user, one model at a timeScales with the serviceConcurrent requests at peak
Data pathNothing leaves the machineRequests go to your deploymentData classification rules
Cost shapeHardware, power and timeFlat monthly plansUtilisation and budget

Frequently asked questions

Can the RTX 4060 run a 13B model?

Only with aggressive quantization and partial offload. A 13B model at 4-bit is roughly 7-8 GB before KV cache, so it will not fit comfortably in 8 GB and layers will spill to system RAM.

Why do I get out-of-memory errors with a 7B model?

The weights fit, but the KV cache grows with context length and concurrency. Lower the context window, close other GPU apps or reduce batch size.

Is CUDA required?

No, llama.cpp can run on CPU, but CUDA is what makes the GPU fast. Install current NVIDIA drivers and use a runtime built with CUDA support.

How much context can 8 GB handle?

It depends on the model's layers and KV head configuration. Start at 4K, raise it in steps and watch VRAM; the KV cache, not the weights, is usually the limit.

Should I quantize the KV cache?

On runtimes that support it, KV cache quantization extends usable context at some quality cost. Test it against your eval before adopting it.

When should I stop optimising locally?

When the model you need will not fit, when you need concurrency, or when long prompts dominate. A hosted OpenAI-compatible endpoint covers those cases without a code rewrite.

Does the 4060 work for local RAG?

Yes for small corpora and modest context. Embeddings are cheap, but retrieval-augmented prompts are long, so budget the KV cache carefully.