Local AI

How do you serve a local AI API?

A local AI API is an HTTP server that loads model weights and exposes inference over a network, usually with OpenAI-compatible routes. llama.cpp server, Ollama, LM Studio and vLLM all serve /v1/chat/completions locally. The main design work is not the endpoint itself but authentication, concurrency, model loading strategy and how you route overflow to a hosted API.

Key facts

Common serversllama.cpp server, Ollama, LM Studio and vLLM
Wire formatOpenAI-compatible /v1/chat/completions and /v1/embeddings on most servers
Default portsOllama 11434, LM Studio 1234, vLLM 8000, llama.cpp configurable
ConcurrencyvLLM uses continuous batching; llama.cpp and Ollama queue or batch differently
AuthLocal servers often ship without authentication and must be bound to localhost or placed behind a proxy
Hybrid optionPlugsky is OpenAI-compatible with 30+ models, so the same client can switch base URL
Endpoint statusChat, streaming, JSON mode, function calling, embeddings, RAG and agents are live

TL;DR

  • Every major local runtime speaks enough of the OpenAI API to reuse existing SDKs.
  • Bind local servers to localhost by default; add a reverse proxy with keys before exposing them.
  • Choose vLLM for concurrent traffic, llama.cpp or Ollama for single-user and CPU-friendly setups.
  • Model loading strategy and context size dominate latency, not just raw tokens per second.
  • Keep the endpoint OpenAI-compatible so overflow can route to a hosted API without code changes.

How it works, step by step

  1. Pick a runtime based on load: vLLM for concurrent GPU serving, Ollama or llama.cpp for simple or CPU-first use.
  2. Load a quantized model and confirm the /v1/models route lists it.
  3. Send a chat completion with curl and verify streaming works before integrating an SDK.
  4. Add authentication: bind to 127.0.0.1 or put an API-key-checking reverse proxy in front.
  5. Set model lifetime and context limits so idle models unload and memory stays predictable.
  6. Add metrics and request logs, then load-test with your real concurrency.
  7. Configure a fallback route to an OpenAI-compatible cloud API for spikes and failures.
1Pick a runtimebased on load: vLLMfor concurrent GPU2Load a quantizedmodel and confirmthe /v1/models3Send a chatcompletion withcurl and verify4Add authentication:bind to 127.0.0.1or put an5Set model lifetimeand context limitsso idle models6Add metrics andrequest logs, thenload-test with your

Original data

OpenAI-compatiWire formatOllama 11434, Default portsPlugsky is OpeHybrid optionSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the OpenAI-compatible API tester →

What a local AI API includes

At minimum, a local API server does four things: load weights into memory, tokenize prompts, run inference, and return responses in a documented shape. Most runtimes add model listing, streaming, stop sequences and parameter passthrough. The OpenAI-compatible shape matters because it lets you keep existing SDKs, evaluation scripts and agent frameworks unchanged.

llama.cpp server is the most portable, running on CPU, Metal, CUDA, ROCm and Vulkan. Ollama wraps the same family of runtimes with a model registry and a simple daemon. vLLM targets GPU throughput with paged attention and continuous batching. LM Studio packages a GUI plus a server for desktop workflows.

Concurrency, memory and model loading

Single-user servers feel simple until two requests arrive. llama.cpp and Ollama can process requests serially or with limited parallel slots, so latency grows with queue depth. vLLM batches concurrent requests on the GPU, which raises throughput substantially for many short generations.

Memory is the other half. Weights are fixed, but the KV cache grows with context length and concurrency. Long contexts and parallel requests multiply cache usage, so cap context size, limit concurrent sequences, and monitor memory before adding another model. Loading several models at once is rarely worth it on a single GPU; route by model with a queue or a second server instead.

Exposure, security and hybrid routing

Local does not mean exposed by accident. Default installs often listen on all interfaces with no authentication, which turns a workstation into an open inference endpoint on the LAN. Bind to localhost, or place the server behind a reverse proxy that checks API keys, enforces rate limits and terminates TLS. Log requests without storing full prompts unless you need them, and separate keys per application.

For capacity beyond one machine, keep the local API as the first hop and add a hosted fallback. Plugsky exposes an OpenAI-compatible API with 30+ models on one key, and chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live. Audio, image, moderation, batch and fine-tuning endpoints are coming soon, so keep those on their current provider. Compare plans on the live pricing page.

Honest comparison

ServerBest forConcurrencyHardware reachOpenAI-compatible
llama.cpp serverControl and portabilityLimited parallel slotsCPU, Metal, CUDA, ROCm, VulkanYes
OllamaFast local startsQueue with limited parallelismCPU, Metal, CUDA, ROCmYes
LM StudioDesktop evaluationLowCPU, Metal, CUDAYes
vLLMConcurrent GPU servingContinuous batchingCUDA, ROCm, some CPUYes
PlugskyTeam and overflow capacityManagedHostedYes

Frequently asked questions

Which local server should I start with?

Ollama or LM Studio for a quick start, llama.cpp server when you want quantization and offload control, and vLLM when multiple users hit the endpoint at once.

Is the local API really OpenAI-compatible?

Most servers implement the chat completions, streaming and models routes. Feature coverage for tools, JSON mode and embeddings varies, so test the exact routes you depend on.

How do I secure a local inference server?

Bind it to 127.0.0.1, or put a reverse proxy in front that enforces API keys, rate limits and TLS. Never expose an unauthenticated inference port to the internet.

Can I run a local API without a GPU?

Yes. llama.cpp on CPU works for small quantized models, though throughput suits batch or occasional use rather than heavy interactive traffic.

How do I handle more traffic than one machine can serve?

Add a queue, then route overflow to a hosted OpenAI-compatible API. Because the wire format matches, clients do not change.

Does Plugsky offer a local deployment?

Yes. Plugsky supports cloud, VPC, on-prem and air-gapped deployment options for enterprise customers.

What if I need audio or image endpoints?

Plugsky labels audio, image, moderation, batch and fine-tuning endpoints as coming soon. Keep those workloads on their current provider until they ship.