Key facts
| VRAM | 24 GB GDDR6X, the largest pool in the consumer tier |
| Comfortable range | 13B at 8-bit or a 30B-class model at 4-bit with context |
| Serving engines | llama.cpp and Ollama for single user, vLLM for concurrency |
| KV cache | Quantization can extend context at a small quality cost |
| Throughput features | vLLM adds continuous batching and paged attention |
| Multi-GPU | Two cards can split a model with tensor parallelism |
| Cloud fallback | Plugsky serves 30+ models over one OpenAI-compatible API |
| Endpoint status | Chat, streaming, JSON mode, function calling, embeddings, RAG and agents live |
TL;DR
- 24 GB comfortably runs 13B at 8-bit or a 30B-class model at 4-bit.
- Pick the engine by user count: llama.cpp for one, vLLM for many.
- KV cache quantization buys context when memory is tight.
- Tensor parallelism across two GPUs handles models that exceed one card.
- Go private cloud when you need multi-tenant scale or 70B-class models.
How it works, step by step
- Install current NVIDIA drivers and the CUDA toolkit your runtime expects.
- Choose a target: 13B at 8-bit for quality, a 30B-class model at 4-bit for capability.
- Set context length from real prompts and monitor VRAM headroom.
- Switch to vLLM with continuous batching if more than one user hits the model.
- Quantize the KV cache if you need long context on a tight budget.
- Measure throughput at realistic concurrency and prompt lengths.
- Keep an OpenAI-compatible cloud endpoint for overflow and larger models.
Original data
Try it yourself
What 24 GB unlocks
At 8-bit, a 13B model occupies roughly 13-14 GB of weights, leaving about 10 GB for the KV cache and runtime: enough for long prompts and a comfortable context window. At 4-bit, 30B-class models land near 17-19 GB, still leaving room for moderate context. A 70B model at 4-bit is roughly 35-40 GB, so it will not fit on one card without offload.
That places the 4090 in a useful middle tier: strong single-user capability and, with the right server, credible small-team throughput. Decide early whether the card is a workstation tool or a shared endpoint, because the engine choice follows from that.
Engine choice: llama.cpp, Ollama or vLLM
For one person at a time, llama.cpp-based runtimes are simple and flexible. Ollama wraps that engine with model management and an OpenAI-compatible server; LM Studio adds a GUI. You get easy quantization control, CPU offload fallback and Metal-style portability across hardware.
vLLM is designed for serving: continuous batching, paged attention and an OpenAI-compatible endpoint. It delivers far higher total throughput when requests overlap, which is why it is the usual pick for a small internal API. The trade-offs are Linux-first tooling, GPU-resident memory and less tolerance for partial offload.
- Single user, quick experiments: Ollama or llama.cpp.
- Team-facing endpoint: vLLM behind an OpenAI-compatible gateway.
- Very large models: multiple GPUs with tensor parallelism, or a hosted endpoint.
Serving multiple users without a rewrite
Whichever engine you choose, expose an OpenAI-compatible /v1 surface. That keeps clients identical whether the model runs on the 4090 or in the cloud, and it makes the capacity decision reversible. Add request limits, timeouts and queueing so one long generation does not starve the rest.
Plan for the ceiling. A single 4090 serves a small team on a mid-size model, but concurrent long-context traffic and 70B-class work exceed one card. Plugsky's OpenAI-compatible API is the overflow path: 30+ models, flat monthly self-serve plans, and VPC, on-prem or air-gapped deployment for teams that cannot use a public endpoint. Chat, streaming, JSON mode, tools, embeddings, RAG and agents are live; check pricing for the current plans.
Honest comparison
| Concern | RTX 4090 24 GB | Plugsky private deployment | Check before deciding |
|---|---|---|---|
| Model ceiling | 13B-30B class; 70B needs offload or two cards | 30+ hosted models including frontier tiers | Largest model your workload needs |
| Concurrency | Good with vLLM, bounded by one card | Scales with the deployment | Peak concurrent requests |
| Data control | Local by construction | VPC, on-prem and air-gapped options | Compliance requirements |
| Operations | Drivers, CUDA, engines and monitoring | Managed, with SLA and status page | Team capacity |
| Cost shape | Hardware, power and cooling | Flat monthly plans | Utilisation and growth curve |
Frequently asked questions
Can a 4090 run a 70B model?
Not comfortably. A 70B model at 4-bit is roughly 35-40 GB of weights, above 24 GB, so it needs CPU offload, a second GPU or a hosted endpoint.
Should I use vLLM or llama.cpp on a 4090?
Use llama.cpp or Ollama for a single user and quick experiments. Use vLLM when requests arrive concurrently, since continuous batching and paged attention raise total throughput.
Is 8-bit quantization worth it?
Q8 is near-lossless and fits a 13B-class model in 24 GB. Choose it when output quality matters and the model fits; drop to 4-bit only for larger models.
How do I extend context without another GPU?
Cap context to what prompts actually use and enable KV cache quantization where supported. Both trade a little quality for memory.
Can two 4090s serve larger models?
Yes. Tensor parallelism can split a model across two cards, though interconnect and PCIe bandwidth affect scaling. Verify the runtime supports the split for your model family.
Is the 4090 good for fine-tuning?
It is a common entry point for parameter-efficient methods such as LoRA on small and mid-size models. Full fine-tuning of large models needs far more memory and is usually better on hosted training.
When should I move off the 4090?
When you need multi-tenant concurrency, models beyond 24 GB, or guaranteed uptime. An OpenAI-compatible private endpoint removes the single-machine limit.