Key facts
| Families | Instruction-tuned families such as Llama, Qwen, Mistral and gpt-oss |
| Size classes | 3B-4B small, 7B-8B general, 13B-30B capable, 70B+ server class |
| Memory rule | 4-bit weights are roughly half a gigabyte per billion parameters |
| Quantization | Q4_K_M is a common default; Q8 is near-lossless |
| Licences | Terms vary by family and sometimes by size or region |
| Task fit | Tool calling, coding, multilingual and vision needs differ by model |
| Hosted alternative | Plugsky serves 30+ models over one OpenAI-compatible API |
| Endpoint status | Chat, streaming, tools, JSON mode, embeddings, RAG and agents live |
TL;DR
- Start from the task, not from a leaderboard.
- The 7B-8B class is the general-purpose sweet spot.
- At 4-bit, budget roughly half a gigabyte per billion parameters.
- Check the licence before commercial deployment.
- Keep a hosted API for models that will not fit locally.
How it works, step by step
- Write down the task, language coverage and structured-output needs.
- Pick a family known for that task and choose the smallest size that plausibly works.
- Quantize to fit your memory budget, starting at 4-bit.
- Test on your own examples, including failure and edge cases.
- Check the licence for your intended commercial use.
- Pin the exact model revision once it passes.
- Route anything that fails to a hosted OpenAI-compatible model.
Original data
Try it yourself
Open the local model recommender →
How to shortlist a local model
Begin with the task, not the model rankings. Extraction and classification need instruction following, RAG needs reliable grounding and reasonable context length, coding needs code-heavy training data, and multilingual work needs coverage your users actually speak. Each of those points at different families.
Then filter by constraints: hardware memory, required licence, context window and whether tool calling matters. A model that is excellent at writing but emits malformed JSON is a poor choice for an agent pipeline, because every parse failure costs an extra loop.
Size, quantization and memory
Parameter count sets the floor. At 4-bit quantization, allow roughly half a gigabyte per billion parameters for weights, then add the KV cache, runtime overhead and system headroom.
- 3B-4B: fits 8 GB comfortably; good for extraction, classification and light chat.
- 7B-8B: the general-purpose class; comfortable on 16 GB.
- 13B-14B: better reasoning and writing; wants 24 GB at 4-bit, more at higher precision.
- 30B-class: strong capability; fits 24 GB only with careful quantization and modest context.
- 70B+: server or multi-GPU territory.
Context length is the hidden variable, because KV cache grows with it.
Licences and hosted fallback
Open weights are not the same as unrestricted use. Licences differ between families and can include conditions on scale, region or redistribution, and quantized repackagings inherit them. Record the licence and the exact revision for every model you deploy, so legal review is repeatable.
Keep a fallback path for tasks your local model cannot handle. Plugsky serves 30+ models over one OpenAI-compatible API with flat monthly self-serve plans, and offers region selection plus VPC, on-prem and air-gapped deployment. Chat, streaming, tools, JSON mode, embeddings, RAG and agents are live; batch and fine-tuning endpoints are coming soon. See pricing for plans.
Honest comparison
| Model class | Typical fit | Memory at 4-bit | Best for |
|---|---|---|---|
| 3B-4B | 8 GB machines and CPU | Roughly 2-3 GB | Extraction, classification, light chat |
| 7B-8B | 16 GB machines | Roughly 4-5 GB | General chat, RAG, coding help |
| 13B-14B | 24 GB GPUs | Roughly 7-9 GB | Better reasoning and writing |
| 30B-class | 24 GB with care | Roughly 17-19 GB | Strong general capability |
| 70B+ | Multi-GPU or server | 35 GB and up | Frontier-style local quality |
Frequently asked questions
How many parameters do I need?
It depends on the task. Extraction and classification work at 3B-4B, general chat and RAG at 7B-8B, and harder reasoning at 13B and above. Always test on your own examples.
What does 4-bit quantization cost in quality?
It is usually a modest drop on common tasks, but it varies by model and workload. Compare quantized and full-precision outputs on your evaluation set before deciding.
How much memory per parameter?
At 4-bit, allow roughly half a gigabyte per billion parameters for weights, then add KV cache, runtime overhead and system headroom.
Do licences matter for local use?
Yes. Some families restrict commercial use or add conditions, and terms can differ by model size or region. Record the licence with the deployed revision.
Should I choose one model or several?
Keep one general model and add specialists only where they clearly win, because each extra model adds storage, memory and version management.
What about community fine-tunes?
They can be excellent, but provenance, licence and evaluation quality vary. Prefer documented sources and verify before production use.
What if no local model is good enough?
Use a hosted OpenAI-compatible API for those workloads. Plugsky serves 30+ models with flat monthly plans and private deployment options, so the local model can stay your first choice.