Local AI

What does it take to run AI locally from start to finish?

Running AI locally end to end means four layers working together: hardware with enough fast memory, a runtime that serves an OpenAI-compatible API, a quantized model that fits your budget, and an application layer that adds retrieval and tools. Start with one real task, prove quality and latency, then harden security and plan hybrid routing for workloads that outgrow the machine.

Key facts

Layer 1Hardware sized by fast memory: VRAM, unified memory or system RAM
Layer 2Runtime such as Ollama, llama.cpp, LM Studio or vLLM
Layer 3Quantized model: roughly 4-5 GB of weights per 7B model at 4-bit
Layer 4Application speaking OpenAI-compatible /v1 with optional RAG and tools
Memory ruleWeights plus KV cache plus overhead must fit; cache grows with context and concurrency
SecurityBind localhost, authenticate access, sandbox tools and encrypt storage
Hybrid optionPlugsky serves 30+ models over the same OpenAI-compatible interface

TL;DR

  • Build one layer at a time and validate with a real task before adding complexity.
  • Fit the model to memory; quantization is the main lever.
  • Keep the API OpenAI-compatible for portability between runtimes and cloud.
  • Add RAG only when the task needs external knowledge, and tools only when one call is not enough.
  • Design hybrid routing early so scale does not force a rewrite.

How it works, step by step

  1. Pick one workload with clear success criteria and privacy requirements.
  2. Size memory from weights, KV cache and concurrency, then choose a model that fits.
  3. Install a runtime and confirm the OpenAI-compatible endpoint works.
  4. Quantize or select a quantization that holds quality on your evaluation set.
  5. Add retrieval or tools only if the workload requires them.
  6. Secure the deployment and add monitoring for latency, memory and errors.
  7. Define a hybrid policy and route overflow or hard tasks to a hosted API.
1Pick one workloadwith clear successcriteria and2Size memory fromweights, KV cacheand concurrency,3Install a runtimeand confirm theOpenAI-compatible4Quantize or selecta quantization thatholds quality on5Add retrieval ortools only if theworkload requires6Secure thedeployment and addmonitoring for

Original data

Quantized modeLayer 3Application spLayer 4Plugsky servesHybrid optionSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the local model recommender →

Stage 1: choose the workload and size the hardware

Start from a task, not a model. Good first workloads are document summarisation, internal question answering over a fixed corpus, drafting assistance and classification. They have measurable outputs and bounded context. Avoid starting with an autonomous agent that needs many tools, because failures are hard to attribute.

Then size memory. Weights are parameters times bits per weight: a 7B model at 4-bit needs roughly 4-5 GB, a 14B model about 8-9 GB, a 32B model about 17-19 GB. KV cache adds to that and grows with context length and concurrent requests. If the numbers do not fit, choose a smaller model or a higher quantization level rather than accepting constant offload.

Stage 2: pick the runtime and model format

The runtime decides where the model can run. llama.cpp covers CPU, Metal, CUDA, ROCm and Vulkan with GGUF files and tunable quantization; Ollama wraps it with simpler operations; LM Studio packages a GUI and server; vLLM serves GPU traffic with batching. All four expose OpenAI-compatible routes, which keeps client code portable.

  • CPU or laptop: llama.cpp or Ollama with GGUF at Q4-Q5.
  • Apple Silicon: llama.cpp with Metal, or MLX-native formats.
  • GPU server: vLLM with a GPU-native 4-bit format.

Test structured output — JSON mode and tool calls — on the exact quantized checkpoint. Precision loss shows up there before it shows up in fluent prose.

Stage 3: retrieval, tools and security

Add retrieval when answers depend on private documents. Chunk at 256-512 tokens with modest overlap, embed with a model that covers your languages, retrieve broadly and rerank narrowly. Require citations so grounding is checkable. Add tools only when a single model call cannot complete the task, and keep the tool set small with typed schemas.

Security is a design stage, not a final check. Restrict network exposure, authenticate every client, run tools in a sandbox with allowlists, keep secrets out of prompts and logs, and encrypt storage. Treat retrieved documents and tool output as untrusted, because prompt injection travels with content.

Stage 4: production and hybrid routing

Production adds monitoring, evaluation and capacity planning. Track latency percentiles, queue depth, cache memory and error rates, and keep a fixed evaluation suite so model or prompt changes are measured rather than guessed. Plan upgrades: quantized checkpoints, runtime versions and embedding models all affect output, so re-run the suite after each change.

Most teams end up hybrid. Local inference handles private, routine and high-volume work; a hosted API absorbs hard reasoning, larger models and spikes. Because both sides speak the OpenAI-compatible interface, routing is configuration rather than a rewrite. Plugsky serves 30+ models with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live, and supports cloud, VPC, on-prem and air-gapped deployment. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. Review the live pricing page before sizing a plan.

Honest comparison

StageGoalMain decisionSuccess signal
WorkloadPick one real taskTask boundedness and data sensitivityMeasurable output quality
HardwareFit weights and cacheMemory capacity and bandwidthModel runs without constant offload
RuntimeServe an APIConcurrency and hardware reachStable latency under real load
Retrieval and toolsAdd knowledge and actionChunking, embeddings, tool schemasGrounded answers and valid calls
ProductionOperate and scaleMonitoring, evaluation, hybrid policyQuality holds as usage grows

Frequently asked questions

What is the minimum hardware for local AI?

An 8 GB GPU or a machine with 16 GB of unified memory can run 7B-8B models at 4-bit. CPU-only setups work for small models and batch jobs.

Which model should a beginner start with?

A 7B-8B instruction-tuned model at 4-bit is a good first choice. It fits modest hardware and handles summarisation, drafting and simple question answering.

Do I need RAG to run local AI?

No. Many tasks work with a plain chat model. Add retrieval when answers depend on documents or data the model does not know.

How do I know if my local model is good enough?

Define success criteria before starting, then test on 20-50 real prompts. Compare against a hosted model to establish the quality gap you are accepting.

Is local AI cheaper than a hosted API?

It depends on utilization. Local hardware has upfront and operational costs but no per-token fee. Hosted plans make sense for spiky or frontier-model workloads.

How do I keep prompts and documents private?

Keep inference, storage and retrieval local, restrict tool network access, encrypt disks and avoid logging sensitive text. A hybrid route sends only the requests you choose.

When should I move to a hosted API?

When concurrency, context length or model size exceeds local capacity, or when quality on hard tasks falls short. Keep the same OpenAI-compatible interface to make the move cheap.

What are the most common local AI mistakes?

Oversized models that spill to disk, skipping evaluation, exposing inference ports without authentication, and ignoring KV cache growth at long context.