Key facts
| Runtimes | Ollama, llama.cpp server, LM Studio and vLLM expose OpenAI-compatible routes |
| Compatibility | Chat completions, streaming and model listing are the core surface |
| Gateway | Add authentication, rate limits and logging in front of the runtime |
| Keys | Issue per-application keys instead of sharing one |
| Network | Bind to localhost or an internal interface, not the public internet |
| Model naming | Keep a model alias map so clients do not hard-code versions |
| Managed option | Plugsky provides the same interface with 30+ models, managed |
| Endpoint status | Chat, streaming, tools, JSON mode, embeddings, RAG and agents live |
TL;DR
- Self-host by running a runtime with OpenAI-compatible routes.
- Put authentication and rate limits in a gateway in front of it.
- Issue per-application keys and never expose the runtime directly.
- Keep a model alias map so versions are configuration.
- The same client can switch to a managed endpoint by changing the base URL.
How it works, step by step
- Pick a runtime: Ollama, llama.cpp server, LM Studio or vLLM.
- Confirm the routes your clients need: chat, streaming, embeddings.
- Bind the runtime to localhost or an internal interface.
- Add a gateway for authentication, rate limits and request logging.
- Issue per-application keys and a model alias map.
- Test the same client code against local and hosted endpoints.
- Monitor latency, errors and memory, then plan capacity.
Try it yourself
Open the OpenAI-compatible API tester →
Why standardise on an OpenAI-compatible API
The OpenAI request shape has become the common interface for AI clients. Frameworks, SDKs and internal tools already speak it. If your local endpoint speaks it too, you avoid bespoke client code and keep the ability to move between local, private and hosted inference without touching application logic.
It also simplifies comparison. When every candidate exposes the same routes, you can switch models and runtimes with configuration instead of code, which is exactly what you want while evaluating quality and cost.
Building the service safely
A runtime alone is not a service. Put a gateway in front of it and make that gateway the only thing clients talk to.
- Authentication: issue per-application keys and rotate them on a schedule.
- Rate limits: protect the GPU and stop one client from starving others.
- Logging: record requests, latencies and errors for debugging and audit.
- Network: bind to localhost or an internal interface; do not publish the runtime directly.
- Model aliases: let clients request stable names while you control the real version behind them.
Test compatibility deliberately: streaming and tool calling vary slightly between runtimes and can break clients that only tested basic chat.
Scaling and managed alternatives
Self-hosting is a good fit while demand is predictable and one machine can serve it. Concurrency, larger models and uptime expectations are the usual breaking points, because each requires hardware, redundancy and operational attention that a workstation setup does not have.
When that happens, a managed endpoint that exposes the same interface removes the operational work. Plugsky serves 30+ models over one OpenAI-compatible API with flat monthly self-serve plans and region selection plus VPC, on-prem and air-gapped deployment. Chat, streaming, tools, JSON mode, embeddings, RAG and agents are live; audio, image, batch and fine-tuning endpoints are coming soon. See pricing for plans and start free with plugsky-micro and plugsky-lite.
Honest comparison
| Concern | Self-hosted runtime | Plugsky managed endpoint | Check before deciding |
|---|---|---|---|
| Interface | OpenAI-compatible, self-run | OpenAI-compatible, managed | Client needs |
| Auth | You add a gateway | Platform keys and controls | Security requirements |
| Scaling | Your hardware | Service capacity | Peak load |
| Availability | You own uptime | SLA and status page | Criticality |
| Operations | You patch and monitor | Managed | Team capacity |
Frequently asked questions
Which runtime is best for an OpenAI-compatible API?
Ollama for simplicity, llama.cpp server for control, LM Studio for desktop use, and vLLM for concurrent GPU serving. All expose compatible routes.
Do I need a gateway?
For anything beyond local testing, yes. A gateway handles authentication, rate limits, logging and key rotation, which the runtimes do not provide robustly on their own.
Can I expose it on the internet?
Technically yes, but not directly. Terminate TLS and authentication at a gateway, restrict networks and monitor abuse. Internal-only is the safer default.
How do I handle model version changes?
Keep a model alias map in your gateway so clients request stable names while the mapping to real versions is controlled centrally.
Does streaming work?
Yes. Streaming chat completions are part of the compatible surface, though behaviour differs slightly between runtimes, so test with your client.
What about embeddings?
Most runtimes can serve embedding models through compatible routes. Confirm dimensions and keep them consistent with your vector store.
When should I switch to a managed endpoint?
When concurrency, model size, uptime or operations become a burden. Because the API shape is the same, switching is a base URL and key change.