Key facts
| What vLLM is | Open-source GPU inference server |
| Performance design | Paged attention and continuous batching for concurrency |
| API | OpenAI-compatible server |
| Requirements | GPUs plus serving operations |
| Managed model | Plugsky serves 30+ models with no infrastructure to run |
| Pricing | Flat monthly self-serve plans; free plan with two models |
| Deployment | Cloud, VPC, on-prem or air-gapped |
| Status | Chat, streaming, JSON mode, function calling and embeddings live on Plugsky |
TL;DR
- vLLM maximises control and throughput on GPUs you operate.
- A managed cloud removes GPU work and makes cost predictable.
- High sustained utilisation favors vLLM; variable load favors managed.
- Hybrid routing lets you use both without rewriting clients.
- Evaluate on your real concurrency and prompt lengths, not benchmarks.
How it works, step by step
- Estimate peak and average concurrency plus latency targets.
- Calculate GPU capacity needed, including headroom for spikes.
- Compare total cost: hardware, operations, engineering and managed plans.
- Load-test vLLM on target hardware and the managed endpoint with the same prompts.
- Decide routing: steady load self-hosted, burst and new workloads managed.
- Keep both behind one OpenAI-compatible interface and revisit quarterly.
Try it yourself
Open the self-hosting break-even calculator →
What vLLM gives you
vLLM is the default answer for self-hosted GPU inference at scale. Paged attention and continuous batching keep GPUs busy across many simultaneous requests, and the server exposes an OpenAI-compatible API, so application code treats it like any other endpoint. It also gives you control over quantisation, parallelism, sampling and model versions.
That control has a price: GPUs to procure, drivers and engines to maintain, autoscaling to build, failures to page on, and engineers to own all of it. Self-hosting is a product decision as much as an infrastructure one.
What a managed cloud changes
A managed platform removes the entire serving layer. Plugsky runs the models and exposes 30+ of them behind one OpenAI-compatible endpoint, so there are no GPUs to size and no engine to patch. Self-serve plans are flat monthly with unlimited fair use on paid tiers, the free plan includes plugsky-micro and plugsky-lite, and enterprise deployment extends to VPC, on-prem and air-gapped when isolation is required. Current plans are on the live pricing page.
The honest limitation is depth of control. Managed inference does not let you choose kernels, custom quantisations or sampling internals. If those are product requirements rather than preferences, vLLM remains the right foundation.
The economics and the hybrid
Self-hosting wins when utilisation is high and stable, engineering capacity exists, and the workload is well understood. Managed wins when demand is spiky or young, when time to production matters, and when a small team should not carry on-call for inference.
The strongest architecture is often both. Keep a vLLM cluster for steady, high-volume paths where you have optimised the model, and route burst, experimentation and less critical workloads to a managed catalogue. Because both sides speak OpenAI-compatible chat, routing between them is configuration, not a rewrite.
Honest comparison
| Dimension | Plugsky (managed) | vLLM self-hosted | What to weigh |
|---|---|---|---|
| Infrastructure | None to operate | GPUs, drivers, autoscaling | Engineering capacity |
| Control | Catalogue and parameters | Kernels, quantisation, versions | How deep your needs go |
| Cost shape | Flat monthly self-serve plans | Hardware plus operations | Utilisation profile |
| Time to production | Minutes on the free plan | Weeks to build and tune | Roadmap pressure |
| Scaling | Managed by the platform | You size and scale clusters | Spike behaviour |
| Deployment options | Cloud, VPC, on-prem, air-gapped | Anywhere you run GPUs | Residency and sovereignty |
Frequently asked questions
Is vLLM better than a managed API?
Neither is universally better. vLLM offers more control and can be cheaper at high sustained utilisation; a managed API removes operations and is faster to production.
Can I use the OpenAI SDK with vLLM?
Yes. vLLM exposes an OpenAI-compatible server, so the same client works with a different base URL. The same is true of Plugsky.
How do I decide between them?
Estimate utilisation, measure your real concurrency, and compare total cost including engineering time. Low or variable utilisation usually favours a managed service.
Does Plugsky support the same models as vLLM?
Plugsky serves a curated catalogue of 30+ models across families, while vLLM can serve whatever weights you can host. Check the live catalogue against the models you need.
What about latency and throughput?
Benchmark both with your prompts, concurrency and context lengths. Published numbers are directional; your workload and hardware decide the result.
Can I run a hybrid?
Yes. Route steady high-volume traffic to self-hosted vLLM and burst or new workloads to a managed endpoint, behind one OpenAI-compatible interface.
Does Plugsky offer air-gapped deployment?
Yes, as an enterprise option alongside VPC and on-prem, for organisations that cannot use a public endpoint.