Key facts
| llama.cpp | Portable C/C++ inference engine for GGUF models |
| Ollama | Local runner built around llama.cpp with a registry |
| vLLM | GPU serving engine with paged attention and continuous batching |
| Hardware | llama.cpp and Ollama run on laptops; vLLM targets datacentre GPUs |
| API | Ollama and vLLM both expose OpenAI-compatible servers |
| Concurrency | vLLM is built for many simultaneous requests |
| Ops burden | Rises from llama.cpp to Ollama to vLLM |
| Managed option | Plugsky serves 30+ models without GPU operations |
TL;DR
- llama.cpp is the engine; Ollama is the local runtime; vLLM is the production server.
- Portability and CPU-only machines point to llama.cpp.
- Developer convenience and a model registry point to Ollama.
- Concurrency, throughput and GPU efficiency point to vLLM.
- A managed API removes GPU operations entirely when that is the goal.
How it works, step by step
- Estimate concurrent requests and latency targets for the workload.
- Check the hardware you can use: laptop, single GPU or multi-GPU node.
- Prototype on Ollama, then load-test the candidate engine on target hardware.
- Measure time to first token and throughput under realistic concurrency.
- Choose quantisation and context limits around the engine's constraints.
- Compare the operating cost with a managed API before committing to GPUs.
Try it yourself
Open the vLLM launch command generator →
Engine, runtime, server
The three names sit at different layers, which is why comparisons often confuse. llama.cpp is the inference engine: C/C++ code that loads quantised GGUF weights and runs them efficiently on CPUs and GPUs, including edge devices. It is the portability floor of the ecosystem.
Ollama is a runtime experience built on that class of local serving: a model registry, simple commands and a background server that other tools can call. vLLM is a serving system designed for GPUs in a datacentre, focused on batching many requests efficiently rather than making one developer comfortable.
Matching tool to workload
Start with concurrency. One user chatting needs a responsive local runner; hundreds of simultaneous API calls need batching, memory management and parallelism, which is vLLM's design centre. Then consider hardware: llama.cpp and Ollama run acceptably on laptops and small servers, while vLLM expects serious GPUs.
- Edge or embedded: llama.cpp, for minimal dependencies and CPU support.
- Local development: Ollama, for model management and a simple server.
- Self-hosted production: vLLM, for throughput and concurrency.
- No GPU operations: a managed API that handles serving for you.
The managed option
Self-hosting an inference server is a commitment: drivers, CUDA versions, model upgrades, autoscaling, failover and on-call. If those are not your product, a managed endpoint is the shorter path. Plugsky serves 30+ models behind one OpenAI-compatible API with flat monthly self-serve plans, a free plan covering plugsky-micro and plugsky-lite, and private deployment options for teams that need isolation without building a serving team. Current plan details are on the live pricing page.
Where self-hosting still wins: custom kernels, unusual quantisation, strict control over sampling internals or steady utilisation high enough to amortise GPUs. In those cases vLLM is usually the engine to build on.
Honest comparison
| Dimension | llama.cpp | Ollama | vLLM | Plugsky |
|---|---|---|---|---|
| Layer | Inference engine | Local runtime | Production server | Managed API |
| Hardware | CPU or GPU, very portable | Laptop to workstation | Datacentre GPUs | No hardware |
| Best for | Edge and embedded use | Local development and scripts | High-concurrency serving | Teams without GPU ops |
| API server | Minimal examples | OpenAI-compatible | OpenAI-compatible | OpenAI-compatible |
| Operations | Low | Low to medium | High | None |
| Cost shape | Free, your hardware | Free, your hardware | GPUs plus operations | Flat monthly plans |
Frequently asked questions
Is Ollama built on llama.cpp?
Ollama builds on the llama.cpp family of local inference technology, adding a model registry, simpler commands and a background server on top.
Can vLLM run on a laptop?
It can run on a single GPU in principle, but it is designed for datacentre GPUs and high concurrency. For laptop use, llama.cpp or Ollama is a better fit.
Which is best for a production API?
If you self-host, vLLM is the usual choice for throughput. If you would rather not operate GPUs, a managed OpenAI-compatible API replaces the whole layer.
Can I use the OpenAI SDK with these engines?
Ollama and vLLM both expose OpenAI-compatible servers, so the same client code works with a different base URL. llama.cpp is an engine you embed rather than a full API platform.
Do I need a GPU for llama.cpp?
No. It runs on CPUs, which is part of its appeal for edge and embedded deployments, though throughput is lower than GPU serving.
How do I avoid GPU operations entirely?
Use a managed API. Plugsky serves 30+ models over one OpenAI-compatible endpoint with flat monthly plans and no infrastructure for you to run.
Which handles the most concurrent requests?
vLLM, by design. Paged attention and continuous batching exist specifically to keep GPUs busy across many simultaneous requests.