Key facts
| Common servers | llama.cpp server, Ollama, LM Studio and vLLM |
| Wire format | OpenAI-compatible /v1/chat/completions and /v1/embeddings on most servers |
| Default ports | Ollama 11434, LM Studio 1234, vLLM 8000, llama.cpp configurable |
| Concurrency | vLLM uses continuous batching; llama.cpp and Ollama queue or batch differently |
| Auth | Local servers often ship without authentication and must be bound to localhost or placed behind a proxy |
| Hybrid option | Plugsky is OpenAI-compatible with 30+ models, so the same client can switch base URL |
| Endpoint status | Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live |
TL;DR
- Every major local runtime speaks enough of the OpenAI API to reuse existing SDKs.
- Bind local servers to localhost by default; add a reverse proxy with keys before exposing them.
- Choose vLLM for concurrent traffic, llama.cpp or Ollama for single-user and CPU-friendly setups.
- Model loading strategy and context size dominate latency, not just raw tokens per second.
- Keep the endpoint OpenAI-compatible so overflow can route to a hosted API without code changes.
How it works, step by step
- Pick a runtime based on load: vLLM for concurrent GPU serving, Ollama or llama.cpp for simple or CPU-first use.
- Load a quantized model and confirm the /v1/models route lists it.
- Send a chat completion with curl and verify streaming works before integrating an SDK.
- Add authentication: bind to 127.0.0.1 or put an API-key-checking reverse proxy in front.
- Set model lifetime and context limits so idle models unload and memory stays predictable.
- Add metrics and request logs, then load-test with your real concurrency.
- Configure a fallback route to an OpenAI-compatible cloud API for spikes and failures.
Original data
Try it yourself
Open the OpenAI-compatible API tester →
What a local AI API includes
At minimum, a local API server does four things: load weights into memory, tokenize prompts, run inference, and return responses in a documented shape. Most runtimes add model listing, streaming, stop sequences and parameter passthrough. The OpenAI-compatible shape matters because it lets you keep existing SDKs, evaluation scripts and agent frameworks unchanged.
llama.cpp server is the most portable, running on CPU, Metal, CUDA, ROCm and Vulkan. Ollama wraps the same family of runtimes with a model registry and a simple daemon. vLLM targets GPU throughput with paged attention and continuous batching. LM Studio packages a GUI plus a server for desktop workflows.
Concurrency, memory and model loading
Single-user servers feel simple until two requests arrive. llama.cpp and Ollama can process requests serially or with limited parallel slots, so latency grows with queue depth. vLLM batches concurrent requests on the GPU, which raises throughput substantially for many short generations.
Memory is the other half. Weights are fixed, but the KV cache grows with context length and concurrency. Long contexts and parallel requests multiply cache usage, so cap context size, limit concurrent sequences, and monitor memory before adding another model. Loading several models at once is rarely worth it on a single GPU; route by model with a queue or a second server instead.
Exposure, security and hybrid routing
Local does not mean exposed by accident. Default installs often listen on all interfaces with no authentication, which turns a workstation into an open inference endpoint on the LAN. Bind to localhost, or place the server behind a reverse proxy that checks API keys, enforces rate limits and terminates TLS. Log requests without storing full prompts unless you need them, and separate keys per application.
For capacity beyond one machine, keep the local API as the first hop and add a hosted fallback. Plugsky exposes an OpenAI-compatible API with 30+ models on one key, and chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, image, moderation, batch and fine-tuning endpoints are coming soon, so keep those on their current provider. Compare plans on the live pricing page.
Honest comparison
| Server | Best for | Concurrency | Hardware reach | OpenAI-compatible |
|---|---|---|---|---|
| llama.cpp server | Control and portability | Limited parallel slots | CPU, Metal, CUDA, ROCm, Vulkan | Yes |
| Ollama | Fast local starts | Queue with limited parallelism | CPU, Metal, CUDA, ROCm | Yes |
| LM Studio | Desktop evaluation | Low | CPU, Metal, CUDA | Yes |
| vLLM | Concurrent GPU serving | Continuous batching | CUDA, ROCm, some CPU | Yes |
| Plugsky | Team and overflow capacity | Managed | Hosted | Yes |
Frequently asked questions
Which local server should I start with?
Ollama or LM Studio for a quick start, llama.cpp server when you want quantization and offload control, and vLLM when multiple users hit the endpoint at once.
Is the local API really OpenAI-compatible?
Most servers implement the chat completions, streaming and models routes. Feature coverage for tools, JSON mode and embeddings varies, so test the exact routes you depend on.
How do I secure a local inference server?
Bind it to 127.0.0.1, or put a reverse proxy in front that enforces API keys, rate limits and TLS. Never expose an unauthenticated inference port to the internet.
Can I run a local API without a GPU?
Yes. llama.cpp on CPU works for small quantized models, though throughput suits batch or occasional use rather than heavy interactive traffic.
How do I handle more traffic than one machine can serve?
Add a queue, then route overflow to a hosted OpenAI-compatible API. Because the wire format matches, clients do not change.
Does Plugsky offer a local deployment?
Yes. Plugsky supports cloud, VPC, on-prem and air-gapped deployment options for enterprise customers.
What if I need audio or image endpoints?
Plugsky labels audio, image, moderation, batch and fine-tuning endpoints as coming soon. Keep those workloads on their current provider until they ship.