Local AI

How do you set up a local AI API server?

A local AI API server loads model weights and exposes inference over HTTP, usually through OpenAI-compatible routes. Choose the runtime by workload: llama.cpp for portability, Ollama for simple local use, vLLM for concurrent GPU traffic. Then handle the operational parts — authentication, model loading, context limits, batching and monitoring — and add a hosted fallback for spikes.

Key facts

Runtime choicesllama.cpp server, Ollama, LM Studio and vLLM
InterfaceOpenAI-compatible /v1 routes on all four options
ThroughputvLLM batches concurrent requests; llama.cpp and Ollama have limited parallel slots
MemoryWeights plus KV cache times concurrent sequences
SecurityBind to localhost or front the server with authentication and TLS
Hybrid optionPlugsky serves 30+ models behind an OpenAI-compatible API
Endpoint statusChat, streaming, JSON mode, function calling, embeddings, RAG and agents are live

TL;DR

  • Pick the runtime from your concurrency profile, not from benchmarks alone.
  • Model loading strategy and context limits decide memory stability.
  • Never expose an inference port without authentication and TLS.
  • Monitor queue depth and cache memory, not just tokens per second.
  • Keep the endpoint OpenAI-compatible so clients can fail over to a hosted API.

How it works, step by step

  1. Estimate concurrency and context length for your workload.
  2. Choose the runtime: llama.cpp for portability, Ollama for simplicity, vLLM for throughput.
  3. Load a quantized model and verify the /v1/models route lists it.
  4. Set context and concurrency limits so memory stays within budget.
  5. Add authentication with a reverse proxy, enforce rate limits and terminate TLS.
  6. Instrument request logs, latency, queue depth and memory use.
  7. Configure a fallback route to a hosted OpenAI-compatible API for overflow and failures.
1Estimateconcurrency andcontext length for2Choose the runtime:llama.cpp forportability, Ollama3Load a quantizedmodel and verifythe /v1/models4Set context andconcurrency limitsso memory stays5Add authenticationwith a reverseproxy, enforce rate6Instrument requestlogs, latency,queue depth and

Try it yourself

Open the vLLM launch command generator →

Choosing the serving runtime

Four options cover most needs. llama.cpp server runs nearly anywhere, including CPU, Metal, CUDA, ROCm and Vulkan, and gives explicit control over quantization and layer offload. Ollama wraps the same family of runtimes with a daemon and simple model management, which suits workstations and small internal tools. LM Studio is a desktop app with a capable local server, useful for evaluation. vLLM targets GPU throughput using paged attention, continuous batching and prefix caching.

Concurrency is the deciding factor. A single-user workflow is fine on llama.cpp or Ollama. A shared service with many short requests needs batching, and vLLM's throughput advantage grows with load. Whichever you choose, keep the OpenAI-compatible surface so clients do not depend on runtime-specific behaviour.

Memory, loading and capacity planning

Model weights are the fixed cost; KV cache is the variable one. Cache grows with prompt length, output length and the number of concurrent sequences, so throughput and memory are directly linked. Set a maximum context, cap concurrent sequences and monitor cache usage before increasing either. Multi-model servers sound attractive but divide GPU memory; serving one model well usually beats serving three poorly.

  • Preload the model at startup for predictable latency.
  • Set an idle unload policy if the machine is shared.
  • Cap output tokens to prevent runaway generations.
  • Use prefix caching when many requests share a system prompt.

Security, observability and hybrid routing

Default installs often listen broadly with no authentication. Bind to localhost for single-machine use, or place the server behind a reverse proxy that validates API keys, applies per-key rate limits and terminates TLS. Keep requests logged without storing sensitive prompt text by default, and rotate keys when access changes.

Observability should cover latency percentiles, queue depth, cache memory and error rates by client. These four signals tell you whether to change capacity or configuration before users notice. When demand exceeds one machine, use the same OpenAI-compatible interface to burst to a hosted provider. Plugsky serves 30+ models with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. A hybrid route keeps sensitive traffic local and sends the rest upstream. See the live pricing page for plan options.

Honest comparison

RuntimeConcurrencyHardware reachSetup effortBest for
llama.cpp serverLimited parallel slotsCPU, Metal, CUDA, ROCm, VulkanModeratePortable single-node serving
OllamaQueue, limited parallelismCPU, Metal, CUDA, ROCmLowWorkstations and small tools
LM StudioLowCPU, Metal, CUDALowestDesktop evaluation
vLLMContinuous batchingCUDA, ROCmHighMulti-user GPU services
PlugskyManagedHostedNoneOverflow and hosted capacity

Frequently asked questions

Which local server is fastest?

For concurrent GPU traffic, vLLM typically delivers the highest throughput because of continuous batching. For a single user, llama.cpp and Ollama are fast enough and easier to run.

Do I need a reverse proxy?

Yes, if the server is reachable by more than one machine. A proxy handles authentication, TLS, rate limits and logging that local servers usually omit.

How do I keep memory stable under load?

Cap context length and concurrent sequences, preload the model, and monitor cache usage. Memory grows with concurrency even when weights are fixed.

Can one server host several models?

Possibly, but they share GPU memory. Serving one model at a time with a routing layer usually performs better than loading several simultaneously.

How do I prevent the endpoint from being exposed publicly?

Bind to localhost or a private interface, restrict firewall rules, and require authentication at the proxy. Never publish an unauthenticated inference port.

What metrics should I monitor?

Latency percentiles, queue depth, KV cache memory, error rates and tokens per second. Queue depth and cache memory are the leading indicators of trouble.

How do I survive traffic spikes?

Add a queue and a hosted fallback. An OpenAI-compatible API such as Plugsky can absorb overflow without changes to client code.

Can I deploy this on-prem or air-gapped?

Yes. Local runtimes work without internet, and Plugsky supports cloud, VPC, on-prem and air-gapped deployment for enterprise customers.