Comparisons

Ollama vs vLLM vs llama.cpp: which inference engine should you use?

These three are not peers. llama.cpp is the portable inference engine; Ollama wraps that style of local serving in a friendly CLI and model registry; vLLM is a GPU-first server built for high-concurrency production. Choose llama.cpp for portability, Ollama for local development, vLLM for throughput, or a managed API when you would rather not run GPUs at all.

Key facts

llama.cppPortable C/C++ inference engine for GGUF models
OllamaLocal runner built around llama.cpp with a registry
vLLMGPU serving engine with paged attention and continuous batching
Hardwarellama.cpp and Ollama run on laptops; vLLM targets datacentre GPUs
APIOllama and vLLM both expose OpenAI-compatible servers
ConcurrencyvLLM is built for many simultaneous requests
Ops burdenRises from llama.cpp to Ollama to vLLM
Managed optionPlugsky serves 30+ models without GPU operations

TL;DR

  • llama.cpp is the engine; Ollama is the local runtime; vLLM is the production server.
  • Portability and CPU-only machines point to llama.cpp.
  • Developer convenience and a model registry point to Ollama.
  • Concurrency, throughput and GPU efficiency point to vLLM.
  • A managed API removes GPU operations entirely when that is the goal.

How it works, step by step

  1. Estimate concurrent requests and latency targets for the workload.
  2. Check the hardware you can use: laptop, single GPU or multi-GPU node.
  3. Prototype on Ollama, then load-test the candidate engine on target hardware.
  4. Measure time to first token and throughput under realistic concurrency.
  5. Choose quantisation and context limits around the engine's constraints.
  6. Compare the operating cost with a managed API before committing to GPUs.
1Estimate concurrentrequests andlatency targets for2Check the hardwareyou can use:laptop, single GPU3Prototype onOllama, thenload-test the4Measure time tofirst token andthroughput under5Choose quantisationand context limitsaround the engine's6Compare theoperating cost witha managed API

Try it yourself

Open the vLLM launch command generator →

Engine, runtime, server

The three names sit at different layers, which is why comparisons often confuse. llama.cpp is the inference engine: C/C++ code that loads quantised GGUF weights and runs them efficiently on CPUs and GPUs, including edge devices. It is the portability floor of the ecosystem.

Ollama is a runtime experience built on that class of local serving: a model registry, simple commands and a background server that other tools can call. vLLM is a serving system designed for GPUs in a datacentre, focused on batching many requests efficiently rather than making one developer comfortable.

Matching tool to workload

Start with concurrency. One user chatting needs a responsive local runner; hundreds of simultaneous API calls need batching, memory management and parallelism, which is vLLM's design centre. Then consider hardware: llama.cpp and Ollama run acceptably on laptops and small servers, while vLLM expects serious GPUs.

  • Edge or embedded: llama.cpp, for minimal dependencies and CPU support.
  • Local development: Ollama, for model management and a simple server.
  • Self-hosted production: vLLM, for throughput and concurrency.
  • No GPU operations: a managed API that handles serving for you.

The managed option

Self-hosting an inference server is a commitment: drivers, CUDA versions, model upgrades, autoscaling, failover and on-call. If those are not your product, a managed endpoint is the shorter path. Plugsky serves 30+ models behind one OpenAI-compatible API with flat monthly self-serve plans, a free plan covering plugsky-micro and plugsky-lite, and private deployment options for teams that need isolation without building a serving team. Current plan details are on the live pricing page.

Where self-hosting still wins: custom kernels, unusual quantisation, strict control over sampling internals or steady utilisation high enough to amortise GPUs. In those cases vLLM is usually the engine to build on.

Honest comparison

Dimensionllama.cppOllamavLLMPlugsky
LayerInference engineLocal runtimeProduction serverManaged API
HardwareCPU or GPU, very portableLaptop to workstationDatacentre GPUsNo hardware
Best forEdge and embedded useLocal development and scriptsHigh-concurrency servingTeams without GPU ops
API serverMinimal examplesOpenAI-compatibleOpenAI-compatibleOpenAI-compatible
OperationsLowLow to mediumHighNone
Cost shapeFree, your hardwareFree, your hardwareGPUs plus operationsFlat monthly plans

Frequently asked questions

Is Ollama built on llama.cpp?

Ollama builds on the llama.cpp family of local inference technology, adding a model registry, simpler commands and a background server on top.

Can vLLM run on a laptop?

It can run on a single GPU in principle, but it is designed for datacentre GPUs and high concurrency. For laptop use, llama.cpp or Ollama is a better fit.

Which is best for a production API?

If you self-host, vLLM is the usual choice for throughput. If you would rather not operate GPUs, a managed OpenAI-compatible API replaces the whole layer.

Can I use the OpenAI SDK with these engines?

Ollama and vLLM both expose OpenAI-compatible servers, so the same client code works with a different base URL. llama.cpp is an engine you embed rather than a full API platform.

Do I need a GPU for llama.cpp?

No. It runs on CPUs, which is part of its appeal for edge and embedded deployments, though throughput is lower than GPU serving.

How do I avoid GPU operations entirely?

Use a managed API. Plugsky serves 30+ models over one OpenAI-compatible endpoint with flat monthly plans and no infrastructure for you to run.

Which handles the most concurrent requests?

vLLM, by design. Paged attention and continuous batching exist specifically to keep GPUs busy across many simultaneous requests.