Local AI

How do you choose a local AI model?

Choose by task first, then size, then licence. Small 3B-4B models handle extraction and simple chat, 7B-8B models are the general-purpose default, and 13B-30B models need more memory for better reasoning. Quantization lets a larger model fit, and the licence decides whether you can use it commercially.

Key facts

FamiliesInstruction-tuned families such as Llama, Qwen, Mistral and gpt-oss
Size classes3B-4B small, 7B-8B general, 13B-30B capable, 70B+ server class
Memory rule4-bit weights are roughly half a gigabyte per billion parameters
QuantizationQ4_K_M is a common default; Q8 is near-lossless
LicencesTerms vary by family and sometimes by size or region
Task fitTool calling, coding, multilingual and vision needs differ by model
Hosted alternativePlugsky serves 30+ models over one OpenAI-compatible API
Endpoint statusChat, streaming, tools, JSON mode, embeddings, RAG and agents live

TL;DR

  • Start from the task, not from a leaderboard.
  • The 7B-8B class is the general-purpose sweet spot.
  • At 4-bit, budget roughly half a gigabyte per billion parameters.
  • Check the licence before commercial deployment.
  • Keep a hosted API for models that will not fit locally.

How it works, step by step

  1. Write down the task, language coverage and structured-output needs.
  2. Pick a family known for that task and choose the smallest size that plausibly works.
  3. Quantize to fit your memory budget, starting at 4-bit.
  4. Test on your own examples, including failure and edge cases.
  5. Check the licence for your intended commercial use.
  6. Pin the exact model revision once it passes.
  7. Route anything that fails to a hosted OpenAI-compatible model.
1Write down thetask, languagecoverage and2Pick a family knownfor that task andchoose the smallest3Quantize to fityour memory budget,starting at 4-bit.4Test on your ownexamples, includingfailure and edge5Check the licencefor your intendedcommercial use.6Pin the exact modelrevision once itpasses.

Original data

3B-4B small, 7Size classes4-bit weights Memory ruleQ4_K_M is a coQuantizationPlugsky servesHosted alternativeSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the local model recommender →

How to shortlist a local model

Begin with the task, not the model rankings. Extraction and classification need instruction following, RAG needs reliable grounding and reasonable context length, coding needs code-heavy training data, and multilingual work needs coverage your users actually speak. Each of those points at different families.

Then filter by constraints: hardware memory, required licence, context window and whether tool calling matters. A model that is excellent at writing but emits malformed JSON is a poor choice for an agent pipeline, because every parse failure costs an extra loop.

Size, quantization and memory

Parameter count sets the floor. At 4-bit quantization, allow roughly half a gigabyte per billion parameters for weights, then add the KV cache, runtime overhead and system headroom.

  • 3B-4B: fits 8 GB comfortably; good for extraction, classification and light chat.
  • 7B-8B: the general-purpose class; comfortable on 16 GB.
  • 13B-14B: better reasoning and writing; wants 24 GB at 4-bit, more at higher precision.
  • 30B-class: strong capability; fits 24 GB only with careful quantization and modest context.
  • 70B+: server or multi-GPU territory.

Context length is the hidden variable, because KV cache grows with it.

Licences and hosted fallback

Open weights are not the same as unrestricted use. Licences differ between families and can include conditions on scale, region or redistribution, and quantized repackagings inherit them. Record the licence and the exact revision for every model you deploy, so legal review is repeatable.

Keep a fallback path for tasks your local model cannot handle. Plugsky serves 30+ models over one OpenAI-compatible API with flat monthly self-serve plans, and offers region selection plus VPC, on-prem and air-gapped deployment. Chat, streaming, tools, JSON mode, embeddings, RAG and agents are live; batch and fine-tuning endpoints are coming soon. See pricing for plans.

Honest comparison

Model classTypical fitMemory at 4-bitBest for
3B-4B8 GB machines and CPURoughly 2-3 GBExtraction, classification, light chat
7B-8B16 GB machinesRoughly 4-5 GBGeneral chat, RAG, coding help
13B-14B24 GB GPUsRoughly 7-9 GBBetter reasoning and writing
30B-class24 GB with careRoughly 17-19 GBStrong general capability
70B+Multi-GPU or server35 GB and upFrontier-style local quality

Frequently asked questions

How many parameters do I need?

It depends on the task. Extraction and classification work at 3B-4B, general chat and RAG at 7B-8B, and harder reasoning at 13B and above. Always test on your own examples.

What does 4-bit quantization cost in quality?

It is usually a modest drop on common tasks, but it varies by model and workload. Compare quantized and full-precision outputs on your evaluation set before deciding.

How much memory per parameter?

At 4-bit, allow roughly half a gigabyte per billion parameters for weights, then add KV cache, runtime overhead and system headroom.

Do licences matter for local use?

Yes. Some families restrict commercial use or add conditions, and terms can differ by model size or region. Record the licence with the deployed revision.

Should I choose one model or several?

Keep one general model and add specialists only where they clearly win, because each extra model adds storage, memory and version management.

What about community fine-tunes?

They can be excellent, but provenance, licence and evaluation quality vary. Prefer documented sources and verify before production use.

What if no local model is good enough?

Use a hosted OpenAI-compatible API for those workloads. Plugsky serves 30+ models with flat monthly plans and private deployment options, so the local model can stay your first choice.