Key facts
| Layer 1 | Hardware sized by fast memory: VRAM, unified memory or system RAM |
| Layer 2 | Runtime such as Ollama, llama.cpp, LM Studio or vLLM |
| Layer 3 | Quantized model: roughly 4-5 GB of weights per 7B model at 4-bit |
| Layer 4 | Application speaking OpenAI-compatible /v1 with optional RAG and tools |
| Memory rule | Weights plus KV cache plus overhead must fit; cache grows with context and concurrency |
| Security | Bind localhost, authenticate access, sandbox tools and encrypt storage |
| Hybrid option | Plugsky serves 30+ models over the same OpenAI-compatible interface |
TL;DR
- Build one layer at a time and validate with a real task before adding complexity.
- Fit the model to memory; quantization is the main lever.
- Keep the API OpenAI-compatible for portability between runtimes and cloud.
- Add RAG only when the task needs external knowledge, and tools only when one call is not enough.
- Design hybrid routing early so scale does not force a rewrite.
How it works, step by step
- Pick one workload with clear success criteria and privacy requirements.
- Size memory from weights, KV cache and concurrency, then choose a model that fits.
- Install a runtime and confirm the OpenAI-compatible endpoint works.
- Quantize or select a quantization that holds quality on your evaluation set.
- Add retrieval or tools only if the workload requires them.
- Secure the deployment and add monitoring for latency, memory and errors.
- Define a hybrid policy and route overflow or hard tasks to a hosted API.
Original data
Try it yourself
Open the local model recommender →
Stage 1: choose the workload and size the hardware
Start from a task, not a model. Good first workloads are document summarisation, internal question answering over a fixed corpus, drafting assistance and classification. They have measurable outputs and bounded context. Avoid starting with an autonomous agent that needs many tools, because failures are hard to attribute.
Then size memory. Weights are parameters times bits per weight: a 7B model at 4-bit needs roughly 4-5 GB, a 14B model about 8-9 GB, a 32B model about 17-19 GB. KV cache adds to that and grows with context length and concurrent requests. If the numbers do not fit, choose a smaller model or a higher quantization level rather than accepting constant offload.
Stage 2: pick the runtime and model format
The runtime decides where the model can run. llama.cpp covers CPU, Metal, CUDA, ROCm and Vulkan with GGUF files and tunable quantization; Ollama wraps it with simpler operations; LM Studio packages a GUI and server; vLLM serves GPU traffic with batching. All four expose OpenAI-compatible routes, which keeps client code portable.
- CPU or laptop: llama.cpp or Ollama with GGUF at Q4-Q5.
- Apple Silicon: llama.cpp with Metal, or MLX-native formats.
- GPU server: vLLM with a GPU-native 4-bit format.
Test structured output — JSON mode and tool calls — on the exact quantized checkpoint. Precision loss shows up there before it shows up in fluent prose.
Stage 3: retrieval, tools and security
Add retrieval when answers depend on private documents. Chunk at 256-512 tokens with modest overlap, embed with a model that covers your languages, retrieve broadly and rerank narrowly. Require citations so grounding is checkable. Add tools only when a single model call cannot complete the task, and keep the tool set small with typed schemas.
Security is a design stage, not a final check. Restrict network exposure, authenticate every client, run tools in a sandbox with allowlists, keep secrets out of prompts and logs, and encrypt storage. Treat retrieved documents and tool output as untrusted, because prompt injection travels with content.
Stage 4: production and hybrid routing
Production adds monitoring, evaluation and capacity planning. Track latency percentiles, queue depth, cache memory and error rates, and keep a fixed evaluation suite so model or prompt changes are measured rather than guessed. Plan upgrades: quantized checkpoints, runtime versions and embedding models all affect output, so re-run the suite after each change.
Most teams end up hybrid. Local inference handles private, routine and high-volume work; a hosted API absorbs hard reasoning, larger models and spikes. Because both sides speak the OpenAI-compatible interface, routing is configuration rather than a rewrite. Plugsky serves 30+ models with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live, and supports cloud, VPC, on-prem and air-gapped deployment. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. Review the live pricing page before sizing a plan.
Honest comparison
| Stage | Goal | Main decision | Success signal |
|---|---|---|---|
| Workload | Pick one real task | Task boundedness and data sensitivity | Measurable output quality |
| Hardware | Fit weights and cache | Memory capacity and bandwidth | Model runs without constant offload |
| Runtime | Serve an API | Concurrency and hardware reach | Stable latency under real load |
| Retrieval and tools | Add knowledge and action | Chunking, embeddings, tool schemas | Grounded answers and valid calls |
| Production | Operate and scale | Monitoring, evaluation, hybrid policy | Quality holds as usage grows |
Frequently asked questions
What is the minimum hardware for local AI?
An 8 GB GPU or a machine with 16 GB of unified memory can run 7B-8B models at 4-bit. CPU-only setups work for small models and batch jobs.
Which model should a beginner start with?
A 7B-8B instruction-tuned model at 4-bit is a good first choice. It fits modest hardware and handles summarisation, drafting and simple question answering.
Do I need RAG to run local AI?
No. Many tasks work with a plain chat model. Add retrieval when answers depend on documents or data the model does not know.
How do I know if my local model is good enough?
Define success criteria before starting, then test on 20-50 real prompts. Compare against a hosted model to establish the quality gap you are accepting.
Is local AI cheaper than a hosted API?
It depends on utilization. Local hardware has upfront and operational costs but no per-token fee. Hosted plans make sense for spiky or frontier-model workloads.
How do I keep prompts and documents private?
Keep inference, storage and retrieval local, restrict tool network access, encrypt disks and avoid logging sensitive text. A hybrid route sends only the requests you choose.
When should I move to a hosted API?
When concurrency, context length or model size exceeds local capacity, or when quality on hard tasks falls short. Keep the same OpenAI-compatible interface to make the move cheap.
What are the most common local AI mistakes?
Oversized models that spill to disk, skipping evaluation, exposing inference ports without authentication, and ignoring KV cache growth at long context.