Local AI

How do you run local AI on an RTX 4090?

The RTX 4090's 24 GB of VRAM runs 13B models at 8-bit or 30B-class models at 4-bit with room for meaningful context. Use llama.cpp or Ollama for single-user work and vLLM for concurrent serving, and quantize the KV cache when long context matters. For multi-user or very large models, an OpenAI-compatible private endpoint is the cleaner path.

Key facts

VRAM24 GB GDDR6X, the largest pool in the consumer tier
Comfortable range13B at 8-bit or a 30B-class model at 4-bit with context
Serving enginesllama.cpp and Ollama for single user, vLLM for concurrency
KV cacheQuantization can extend context at a small quality cost
Throughput featuresvLLM adds continuous batching and paged attention
Multi-GPUTwo cards can split a model with tensor parallelism
Cloud fallbackPlugsky serves 30+ models over one OpenAI-compatible API
Endpoint statusChat, streaming, JSON mode, function calling, embeddings, RAG and agents live

TL;DR

  • 24 GB comfortably runs 13B at 8-bit or a 30B-class model at 4-bit.
  • Pick the engine by user count: llama.cpp for one, vLLM for many.
  • KV cache quantization buys context when memory is tight.
  • Tensor parallelism across two GPUs handles models that exceed one card.
  • Go private cloud when you need multi-tenant scale or 70B-class models.

How it works, step by step

  1. Install current NVIDIA drivers and the CUDA toolkit your runtime expects.
  2. Choose a target: 13B at 8-bit for quality, a 30B-class model at 4-bit for capability.
  3. Set context length from real prompts and monitor VRAM headroom.
  4. Switch to vLLM with continuous batching if more than one user hits the model.
  5. Quantize the KV cache if you need long context on a tight budget.
  6. Measure throughput at realistic concurrency and prompt lengths.
  7. Keep an OpenAI-compatible cloud endpoint for overflow and larger models.
1Install currentNVIDIA drivers andthe CUDA toolkit2Choose a target:13B at 8-bit forquality, a3Set context lengthfrom real promptsand monitor VRAM4Switch to vLLM withcontinuous batchingif more than one5Quantize the KVcache if you needlong context on a6Measure throughputat realisticconcurrency and

Original data

24 GB GDDR6X, VRAM13B at 8-bit oComfortable rangePlugsky servesCloud fallbackSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the VRAM calculator →

What 24 GB unlocks

At 8-bit, a 13B model occupies roughly 13-14 GB of weights, leaving about 10 GB for the KV cache and runtime: enough for long prompts and a comfortable context window. At 4-bit, 30B-class models land near 17-19 GB, still leaving room for moderate context. A 70B model at 4-bit is roughly 35-40 GB, so it will not fit on one card without offload.

That places the 4090 in a useful middle tier: strong single-user capability and, with the right server, credible small-team throughput. Decide early whether the card is a workstation tool or a shared endpoint, because the engine choice follows from that.

Engine choice: llama.cpp, Ollama or vLLM

For one person at a time, llama.cpp-based runtimes are simple and flexible. Ollama wraps that engine with model management and an OpenAI-compatible server; LM Studio adds a GUI. You get easy quantization control, CPU offload fallback and Metal-style portability across hardware.

vLLM is designed for serving: continuous batching, paged attention and an OpenAI-compatible endpoint. It delivers far higher total throughput when requests overlap, which is why it is the usual pick for a small internal API. The trade-offs are Linux-first tooling, GPU-resident memory and less tolerance for partial offload.

  • Single user, quick experiments: Ollama or llama.cpp.
  • Team-facing endpoint: vLLM behind an OpenAI-compatible gateway.
  • Very large models: multiple GPUs with tensor parallelism, or a hosted endpoint.

Serving multiple users without a rewrite

Whichever engine you choose, expose an OpenAI-compatible /v1 surface. That keeps clients identical whether the model runs on the 4090 or in the cloud, and it makes the capacity decision reversible. Add request limits, timeouts and queueing so one long generation does not starve the rest.

Plan for the ceiling. A single 4090 serves a small team on a mid-size model, but concurrent long-context traffic and 70B-class work exceed one card. Plugsky's OpenAI-compatible API is the overflow path: 30+ models, flat monthly self-serve plans, and VPC, on-prem or air-gapped deployment for teams that cannot use a public endpoint. Chat, streaming, JSON mode, tools, embeddings, RAG and agents are live; check pricing for the current plans.

Honest comparison

ConcernRTX 4090 24 GBPlugsky private deploymentCheck before deciding
Model ceiling13B-30B class; 70B needs offload or two cards30+ hosted models including frontier tiersLargest model your workload needs
ConcurrencyGood with vLLM, bounded by one cardScales with the deploymentPeak concurrent requests
Data controlLocal by constructionVPC, on-prem and air-gapped optionsCompliance requirements
OperationsDrivers, CUDA, engines and monitoringManaged, with SLA and status pageTeam capacity
Cost shapeHardware, power and coolingFlat monthly plansUtilisation and growth curve

Frequently asked questions

Can a 4090 run a 70B model?

Not comfortably. A 70B model at 4-bit is roughly 35-40 GB of weights, above 24 GB, so it needs CPU offload, a second GPU or a hosted endpoint.

Should I use vLLM or llama.cpp on a 4090?

Use llama.cpp or Ollama for a single user and quick experiments. Use vLLM when requests arrive concurrently, since continuous batching and paged attention raise total throughput.

Is 8-bit quantization worth it?

Q8 is near-lossless and fits a 13B-class model in 24 GB. Choose it when output quality matters and the model fits; drop to 4-bit only for larger models.

How do I extend context without another GPU?

Cap context to what prompts actually use and enable KV cache quantization where supported. Both trade a little quality for memory.

Can two 4090s serve larger models?

Yes. Tensor parallelism can split a model across two cards, though interconnect and PCIe bandwidth affect scaling. Verify the runtime supports the split for your model family.

Is the 4090 good for fine-tuning?

It is a common entry point for parameter-efficient methods such as LoRA on small and mid-size models. Full fine-tuning of large models needs far more memory and is usually better on hosted training.

When should I move off the 4090?

When you need multi-tenant concurrency, models beyond 24 GB, or guaranteed uptime. An OpenAI-compatible private endpoint removes the single-machine limit.