Local AI

How do you run AI locally, and when should you?

Running AI locally means hosting model inference on hardware you control. The stack has four layers: hardware with enough fast memory, a runtime such as Ollama or llama.cpp, a quantized model that fits your budget, and your application speaking an OpenAI-compatible API. Local AI suits privacy, offline operation and predictable high-volume use; hybrid routing adds hosted capacity when local limits are reached.

Key facts

Stack layersHardware, runtime, model and application interface
Memory ruleWeights plus KV cache plus overhead must fit in fast memory
Common runtimesOllama, llama.cpp, LM Studio and vLLM
Format choiceGGUF for CPU and Apple Silicon; GPU-native formats for server batching
Application interfaceOpenAI-compatible /v1 keeps SDKs and tools unchanged
Hybrid optionPlugsky offers 30+ hosted models behind the same interface
Endpoint statusChat, streaming, JSON mode, function calling, embeddings, RAG and agents are live

TL;DR

  • Fit the model to memory first; speed and precision come after.
  • Start with Ollama or LM Studio, then move to llama.cpp or vLLM as needs grow.
  • Quantize to 4-bit to run larger models on modest hardware.
  • Keep the interface OpenAI-compatible so local and cloud stay swappable.
  • Use hybrid routing for workloads that outgrow one machine.

How it works, step by step

  1. Define the workload: chat, document QA, coding or agents, and its privacy requirements.
  2. Size memory from model weights plus KV cache at your target context length.
  3. Install a runtime and pull a quantized model that fits.
  4. Serve an OpenAI-compatible endpoint and test it with your existing client code.
  5. Add retrieval or tools if the workload requires external knowledge or actions.
  6. Secure the setup: authentication, network exposure, sandboxed tools and logs.
  7. Set a hybrid route for overflow or hard tasks and monitor quality over time.
1Define theworkload: chat,document QA, coding2Size memory frommodel weights plusKV cache at your3Install a runtimeand pull aquantized model4Serve anOpenAI-compatibleendpoint and test5Add retrieval ortools if theworkload requires6Secure the setup:authentication,network exposure,

Try it yourself

Open the local model recommender →

The four layers of a local AI stack

Hardware comes first. Model weights plus KV cache plus runtime overhead must fit in fast memory, whether that is GPU VRAM, unified memory or system RAM. A 7B model at 4-bit needs roughly 4-5 GB before context; a 32B model needs roughly 17-19 GB. Bandwidth then determines how fast tokens appear.

The runtime turns weights into an API. Ollama and LM Studio are the fastest ways to start, llama.cpp gives the most control over quantization and hardware offload, and vLLM delivers throughput for concurrent GPU serving. The model layer is about capability and format: instruction-tuned models for chat and RAG, code models for development, tool-calling models for agents, and quantization levels that match your memory.

The application layer should stay portable. If your client speaks an OpenAI-compatible API, you can move between local runtimes and a hosted provider by changing a base URL and model name, which keeps hybrid routing cheap.

Decisions that matter most

Most local AI projects succeed or fail on three choices. Memory sizing is first, because an oversized model produces constant offload and poor latency. Quantization is second: 4-bit levels such as Q4_K_M are the usual balance, and going lower degrades reasoning and structured output. Interface is third: an OpenAI-compatible endpoint keeps your integration work reusable across runtimes and cloud fallbacks.

  • Match context length to the task and budget KV cache accordingly.
  • Prefer a smaller model that fits over a larger model that spills to disk.
  • Evaluate on your own prompts; public scores rarely predict your workload.
  • Log tool calls and retrieval results, not just final answers.

From first model to production and hybrid

Start small: one runtime, one quantized model, one real task. Prove quality and latency before adding retrieval, tools or multi-user serving. Then harden the setup with authentication, network restrictions, sandboxed tool execution, encrypted storage and request logging. Add retrieval only when the task needs external knowledge, and add agents only when a single model call cannot complete the job.

Production local AI usually becomes hybrid. Local inference handles private, routine or high-volume work; a hosted API absorbs hard reasoning, spikes and workloads that exceed local memory. Plugsky serves 30+ models behind an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live, plus cloud, VPC, on-prem and air-gapped deployment options. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. Review plans on the live pricing page.

Honest comparison

ApproachSetup costPrivacyElasticityBest for
Fully localHardware purchaseHighestLimited to your machinePrivate and offline workloads
Hybrid local plus cloudHardware plus planPolicy-controlledElasticMixed sensitivity and load
Fully hosted APISubscriptionProvider-processedElasticFast starts and spiky traffic
CPU-only localLowestHighVery limitedSmall models and batch jobs

Frequently asked questions

What is local AI in simple terms?

It means the model runs on hardware you control rather than a provider's servers. Prompts and outputs stay on your machine, and you own the runtime, model and updates.

What hardware do I need to start?

An 8 GB GPU or an Apple Silicon machine with 16 GB or more unified memory runs small and mid-size quantized models. CPU-only works for small models and batch use.

Which runtime should I choose first?

Ollama or LM Studio for the quickest start. Move to llama.cpp for quantization and offload control, or vLLM for concurrent GPU serving.

How much does local AI cost?

You pay for hardware, electricity and operations rather than per token. Compare against hosted plans on the live pricing page once you know your real usage pattern.

Is local AI as good as cloud models?

Small local models are strong at routine tasks; frontier hosted models still lead on hard reasoning and very long context. Hybrid routing gets both.

Can local AI work fully offline?

Yes, if the model, tools and data are local. Features that depend on external services, such as web search, will not work without connectivity.

How do I secure a local AI setup?

Authenticate access, avoid exposing inference ports to the internet, sandbox tool execution, encrypt storage and keep logs free of sensitive text where possible.

When should I stop running locally?

When concurrency, context length or model size exceeds your hardware, or when operating costs exceed a hosted plan. An OpenAI-compatible API keeps the move cheap.