Key facts
| Runtime choices | llama.cpp server, Ollama, LM Studio and vLLM |
| Interface | OpenAI-compatible /v1 routes on all four options |
| Throughput | vLLM batches concurrent requests; llama.cpp and Ollama have limited parallel slots |
| Memory | Weights plus KV cache times concurrent sequences |
| Security | Bind to localhost or front the server with authentication and TLS |
| Hybrid option | Plugsky serves 30+ models behind an OpenAI-compatible API |
| Endpoint status | Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live |
TL;DR
- Pick the runtime from your concurrency profile, not from benchmarks alone.
- Model loading strategy and context limits decide memory stability.
- Never expose an inference port without authentication and TLS.
- Monitor queue depth and cache memory, not just tokens per second.
- Keep the endpoint OpenAI-compatible so clients can fail over to a hosted API.
How it works, step by step
- Estimate concurrency and context length for your workload.
- Choose the runtime: llama.cpp for portability, Ollama for simplicity, vLLM for throughput.
- Load a quantized model and verify the /v1/models route lists it.
- Set context and concurrency limits so memory stays within budget.
- Add authentication with a reverse proxy, enforce rate limits and terminate TLS.
- Instrument request logs, latency, queue depth and memory use.
- Configure a fallback route to a hosted OpenAI-compatible API for overflow and failures.
Try it yourself
Open the vLLM launch command generator →
Choosing the serving runtime
Four options cover most needs. llama.cpp server runs nearly anywhere, including CPU, Metal, CUDA, ROCm and Vulkan, and gives explicit control over quantization and layer offload. Ollama wraps the same family of runtimes with a daemon and simple model management, which suits workstations and small internal tools. LM Studio is a desktop app with a capable local server, useful for evaluation. vLLM targets GPU throughput using paged attention, continuous batching and prefix caching.
Concurrency is the deciding factor. A single-user workflow is fine on llama.cpp or Ollama. A shared service with many short requests needs batching, and vLLM's throughput advantage grows with load. Whichever you choose, keep the OpenAI-compatible surface so clients do not depend on runtime-specific behaviour.
Memory, loading and capacity planning
Model weights are the fixed cost; KV cache is the variable one. Cache grows with prompt length, output length and the number of concurrent sequences, so throughput and memory are directly linked. Set a maximum context, cap concurrent sequences and monitor cache usage before increasing either. Multi-model servers sound attractive but divide GPU memory; serving one model well usually beats serving three poorly.
- Preload the model at startup for predictable latency.
- Set an idle unload policy if the machine is shared.
- Cap output tokens to prevent runaway generations.
- Use prefix caching when many requests share a system prompt.
Security, observability and hybrid routing
Default installs often listen broadly with no authentication. Bind to localhost for single-machine use, or place the server behind a reverse proxy that validates API keys, applies per-key rate limits and terminates TLS. Keep requests logged without storing sensitive prompt text by default, and rotate keys when access changes.
Observability should cover latency percentiles, queue depth, cache memory and error rates by client. These four signals tell you whether to change capacity or configuration before users notice. When demand exceeds one machine, use the same OpenAI-compatible interface to burst to a hosted provider. Plugsky serves 30+ models with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. A hybrid route keeps sensitive traffic local and sends the rest upstream. See the live pricing page for plan options.
Honest comparison
| Runtime | Concurrency | Hardware reach | Setup effort | Best for |
|---|---|---|---|---|
| llama.cpp server | Limited parallel slots | CPU, Metal, CUDA, ROCm, Vulkan | Moderate | Portable single-node serving |
| Ollama | Queue, limited parallelism | CPU, Metal, CUDA, ROCm | Low | Workstations and small tools |
| LM Studio | Low | CPU, Metal, CUDA | Lowest | Desktop evaluation |
| vLLM | Continuous batching | CUDA, ROCm | High | Multi-user GPU services |
| Plugsky | Managed | Hosted | None | Overflow and hosted capacity |
Frequently asked questions
Which local server is fastest?
For concurrent GPU traffic, vLLM typically delivers the highest throughput because of continuous batching. For a single user, llama.cpp and Ollama are fast enough and easier to run.
Do I need a reverse proxy?
Yes, if the server is reachable by more than one machine. A proxy handles authentication, TLS, rate limits and logging that local servers usually omit.
How do I keep memory stable under load?
Cap context length and concurrent sequences, preload the model, and monitor cache usage. Memory grows with concurrency even when weights are fixed.
Can one server host several models?
Possibly, but they share GPU memory. Serving one model at a time with a routing layer usually performs better than loading several simultaneously.
How do I prevent the endpoint from being exposed publicly?
Bind to localhost or a private interface, restrict firewall rules, and require authentication at the proxy. Never publish an unauthenticated inference port.
What metrics should I monitor?
Latency percentiles, queue depth, KV cache memory, error rates and tokens per second. Queue depth and cache memory are the leading indicators of trouble.
How do I survive traffic spikes?
Add a queue and a hosted fallback. An OpenAI-compatible API such as Plugsky can absorb overflow without changes to client code.
Can I deploy this on-prem or air-gapped?
Yes. Local runtimes work without internet, and Plugsky supports cloud, VPC, on-prem and air-gapped deployment for enterprise customers.