Key facts
| Runtime | llama.cpp is the reference CPU runtime, with Ollama and LM Studio as easier front ends |
| Bottleneck | Memory bandwidth, not compute, limits tokens per second |
| Vector instructions | AVX2, AVX-512 and AMX on supported CPUs accelerate math |
| Memory | Full model plus cache must fit in system RAM; 16 GB is a floor for small models |
| Quantization | 4-bit levels are standard; more aggressive levels shrink models but reduce quality |
| Best workload | Batch summarisation, classification and offline processing |
| Cloud option | Plugsky serves 30+ models when CPU speed is not enough |
TL;DR
- CPU inference works and needs no GPU, but bandwidth caps generation speed.
- Use 4-bit quantization and small-to-mid models for acceptable results.
- Dual-channel and multi-channel memory configurations matter more than core count.
- It suits batch and offline work better than interactive chat.
- Route latency-sensitive traffic to a GPU or a hosted API.
How it works, step by step
- Check RAM: the model plus cache must fit with room for the operating system.
- Install llama.cpp or a front end such as Ollama and download a 4-bit GGUF model.
- Set the thread count to your physical performance cores, not the logical maximum.
- Enable your CPU's vector instruction set if the build supports it.
- Test tokens per second with a realistic prompt and context length.
- Try a smaller model or a lower quantization level if speed is unusable.
- Offload latency-sensitive work to a GPU machine or a hosted API.
Original data
Try it yourself
Open the self-hosting requirements checker →
How CPU inference works and where time goes
Generating a token requires multiplying activations against the model's weights. On a GPU, those weights sit in very fast memory. On a CPU, they stream from system RAM, and memory bandwidth becomes the ceiling. That is why a CPU with many fast cores can still generate slowly: the cores spend time waiting for data.
Prompt processing is different. It is compute-heavy and parallel, so CPUs can churn through long prompts at a reasonable rate. Generation is the slow, memory-bound part. This asymmetry makes CPU inference acceptable for tasks where output is short relative to input, such as classification or extraction, and painful for long-form generation.
Getting the best from a CPU setup
Quantization is the first lever. A 4-bit model is roughly a quarter of FP16 size and quality is usually adequate for small and mid-size models. Going lower saves memory but degrades reasoning and structured output. Model size is the second lever: a 3B-4B model at 4-bit generates noticeably faster than a 7B-8B model because fewer bytes must stream per token.
- Memory channels: dual-channel or better raises bandwidth substantially.
- Threads: match physical performance cores; oversubscription hurts.
- Instruction sets: ensure your build uses AVX2, AVX-512 or AMX where available.
- NUMA: on multi-socket servers, keep the process on one node.
Expect interactive generation to feel slow on mid-size models and comfortable only on small ones. Batch jobs do not care about tokens per second, which is where CPU inference earns its place.
When CPU is enough and when it is not
CPU-only setups are a good fit for offline summarisation pipelines, document classification, embedding generation at modest volume, and personal use where latency matters less than privacy and simplicity. They are a poor fit for interactive chat with long answers, multi-user services and agent loops that make many model calls per task.
The pragmatic pattern is layered: CPU for offline and low-priority work, a GPU machine when you have one, and a hosted API for traffic that needs speed or model quality beyond the local ceiling. Plugsky serves 30+ models behind an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. The free plan includes plugsky-micro and plugsky-lite, and a 14-day full-access trial covers evaluation. Compare plans on the live pricing page.
Honest comparison
| Setup | Speed profile | Best workload | Memory | Notes |
|---|---|---|---|---|
| CPU, small model at 4-bit | Slow but usable | Occasional chat, tests | 8-16 GB RAM | Simplest to run |
| CPU, mid model at 4-bit | Very slow generation | Batch and offline jobs | 16-32 GB RAM | Throughput over latency |
| CPU embeddings | Moderate | Index building, retrieval | 8-16 GB RAM | Short, parallel calls |
| GPU or cloud API | Fast | Interactive and multi-user | VRAM or hosted | Better for latency |
Frequently asked questions
Can I run a 7B model on CPU?
Yes, at 4-bit quantization with at least 8-16 GB of RAM. Generation will be slow for long answers, so it suits batch or occasional use.
How many tokens per second should I expect?
It varies widely by CPU, memory bandwidth and model size. Small models at 4-bit are usable interactively; mid-size models are usually batch-only on CPU.
Does RAM speed matter?
Yes, a lot. Generation is memory-bandwidth-bound, so faster memory and more channels improve speed more than extra cores.
Is llama.cpp the only option?
It is the most common CPU runtime, and tools like Ollama and LM Studio wrap it. Other engines also offer CPU paths, but llama.cpp is the reference implementation.
Can I serve multiple users from a CPU server?
Not well for interactive use. Throughput and latency degrade quickly; a GPU server or hosted API scales better.
What about CPU embeddings?
Embedding models are small and short-running, so CPU inference works well for building and querying retrieval indexes at moderate volume.
Should I buy a GPU instead?
If interactivity matters, yes. If your work is batch-oriented or privacy-driven and latency is not critical, CPU inference is viable and simpler.
Can I mix CPU and GPU?
Yes. llama.cpp can offload some layers to a GPU while the rest runs on CPU, which speeds up a model that does not fully fit in VRAM.