Local AI

How do you self-host an OpenAI-compatible API?

Run a local runtime that exposes OpenAI-style routes, put it behind your own gateway for authentication and rate limits, and point clients at that base URL. Ollama, llama.cpp server, LM Studio and vLLM all provide compatible endpoints. The goal is one interface your applications use, whether the model runs on a laptop or a GPU server.

Key facts

RuntimesOllama, llama.cpp server, LM Studio and vLLM expose OpenAI-compatible routes
CompatibilityChat completions, streaming and model listing are the core surface
GatewayAdd authentication, rate limits and logging in front of the runtime
KeysIssue per-application keys instead of sharing one
NetworkBind to localhost or an internal interface, not the public internet
Model namingKeep a model alias map so clients do not hard-code versions
Managed optionPlugsky provides the same interface with 30+ models, managed
Endpoint statusChat, streaming, tools, JSON mode, embeddings, RAG and agents live

TL;DR

  • Self-host by running a runtime with OpenAI-compatible routes.
  • Put authentication and rate limits in a gateway in front of it.
  • Issue per-application keys and never expose the runtime directly.
  • Keep a model alias map so versions are configuration.
  • The same client can switch to a managed endpoint by changing the base URL.

How it works, step by step

  1. Pick a runtime: Ollama, llama.cpp server, LM Studio or vLLM.
  2. Confirm the routes your clients need: chat, streaming, embeddings.
  3. Bind the runtime to localhost or an internal interface.
  4. Add a gateway for authentication, rate limits and request logging.
  5. Issue per-application keys and a model alias map.
  6. Test the same client code against local and hosted endpoints.
  7. Monitor latency, errors and memory, then plan capacity.
1Pick a runtime:Ollama, llama.cppserver, LM Studio2Confirm the routesyour clients need:chat, streaming,3Bind the runtime tolocalhost or aninternal interface.4Add a gateway forauthentication,rate limits and5Issueper-applicationkeys and a model6Test the sameclient code againstlocal and hosted

Try it yourself

Open the OpenAI-compatible API tester →

Why standardise on an OpenAI-compatible API

The OpenAI request shape has become the common interface for AI clients. Frameworks, SDKs and internal tools already speak it. If your local endpoint speaks it too, you avoid bespoke client code and keep the ability to move between local, private and hosted inference without touching application logic.

It also simplifies comparison. When every candidate exposes the same routes, you can switch models and runtimes with configuration instead of code, which is exactly what you want while evaluating quality and cost.

Building the service safely

A runtime alone is not a service. Put a gateway in front of it and make that gateway the only thing clients talk to.

  • Authentication: issue per-application keys and rotate them on a schedule.
  • Rate limits: protect the GPU and stop one client from starving others.
  • Logging: record requests, latencies and errors for debugging and audit.
  • Network: bind to localhost or an internal interface; do not publish the runtime directly.
  • Model aliases: let clients request stable names while you control the real version behind them.

Test compatibility deliberately: streaming and tool calling vary slightly between runtimes and can break clients that only tested basic chat.

Scaling and managed alternatives

Self-hosting is a good fit while demand is predictable and one machine can serve it. Concurrency, larger models and uptime expectations are the usual breaking points, because each requires hardware, redundancy and operational attention that a workstation setup does not have.

When that happens, a managed endpoint that exposes the same interface removes the operational work. Plugsky serves 30+ models over one OpenAI-compatible API with flat monthly self-serve plans and region selection plus VPC, on-prem and air-gapped deployment. Chat, streaming, tools, JSON mode, embeddings, RAG and agents are live; audio, image, batch and fine-tuning endpoints are coming soon. See pricing for plans and start free with plugsky-micro and plugsky-lite.

Honest comparison

ConcernSelf-hosted runtimePlugsky managed endpointCheck before deciding
InterfaceOpenAI-compatible, self-runOpenAI-compatible, managedClient needs
AuthYou add a gatewayPlatform keys and controlsSecurity requirements
ScalingYour hardwareService capacityPeak load
AvailabilityYou own uptimeSLA and status pageCriticality
OperationsYou patch and monitorManagedTeam capacity

Frequently asked questions

Which runtime is best for an OpenAI-compatible API?

Ollama for simplicity, llama.cpp server for control, LM Studio for desktop use, and vLLM for concurrent GPU serving. All expose compatible routes.

Do I need a gateway?

For anything beyond local testing, yes. A gateway handles authentication, rate limits, logging and key rotation, which the runtimes do not provide robustly on their own.

Can I expose it on the internet?

Technically yes, but not directly. Terminate TLS and authentication at a gateway, restrict networks and monitor abuse. Internal-only is the safer default.

How do I handle model version changes?

Keep a model alias map in your gateway so clients request stable names while the mapping to real versions is controlled centrally.

Does streaming work?

Yes. Streaming chat completions are part of the compatible surface, though behaviour differs slightly between runtimes, so test with your client.

What about embeddings?

Most runtimes can serve embedding models through compatible routes. Confirm dimensions and keep them consistent with your vector store.

When should I switch to a managed endpoint?

When concurrency, model size, uptime or operations become a burden. Because the API shape is the same, switching is a base URL and key change.