Key facts
| Stack layers | Hardware, runtime, model and application interface |
| Memory rule | Weights plus KV cache plus overhead must fit in fast memory |
| Common runtimes | Ollama, llama.cpp, LM Studio and vLLM |
| Format choice | GGUF for CPU and Apple Silicon; GPU-native formats for server batching |
| Application interface | OpenAI-compatible /v1 keeps SDKs and tools unchanged |
| Hybrid option | Plugsky offers 30+ hosted models behind the same interface |
| Endpoint status | Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live |
TL;DR
- Fit the model to memory first; speed and precision come after.
- Start with Ollama or LM Studio, then move to llama.cpp or vLLM as needs grow.
- Quantize to 4-bit to run larger models on modest hardware.
- Keep the interface OpenAI-compatible so local and cloud stay swappable.
- Use hybrid routing for workloads that outgrow one machine.
How it works, step by step
- Define the workload: chat, document QA, coding or agents, and its privacy requirements.
- Size memory from model weights plus KV cache at your target context length.
- Install a runtime and pull a quantized model that fits.
- Serve an OpenAI-compatible endpoint and test it with your existing client code.
- Add retrieval or tools if the workload requires external knowledge or actions.
- Secure the setup: authentication, network exposure, sandboxed tools and logs.
- Set a hybrid route for overflow or hard tasks and monitor quality over time.
Try it yourself
Open the local model recommender →
The four layers of a local AI stack
Hardware comes first. Model weights plus KV cache plus runtime overhead must fit in fast memory, whether that is GPU VRAM, unified memory or system RAM. A 7B model at 4-bit needs roughly 4-5 GB before context; a 32B model needs roughly 17-19 GB. Bandwidth then determines how fast tokens appear.
The runtime turns weights into an API. Ollama and LM Studio are the fastest ways to start, llama.cpp gives the most control over quantization and hardware offload, and vLLM delivers throughput for concurrent GPU serving. The model layer is about capability and format: instruction-tuned models for chat and RAG, code models for development, tool-calling models for agents, and quantization levels that match your memory.
The application layer should stay portable. If your client speaks an OpenAI-compatible API, you can move between local runtimes and a hosted provider by changing a base URL and model name, which keeps hybrid routing cheap.
Decisions that matter most
Most local AI projects succeed or fail on three choices. Memory sizing is first, because an oversized model produces constant offload and poor latency. Quantization is second: 4-bit levels such as Q4_K_M are the usual balance, and going lower degrades reasoning and structured output. Interface is third: an OpenAI-compatible endpoint keeps your integration work reusable across runtimes and cloud fallbacks.
- Match context length to the task and budget KV cache accordingly.
- Prefer a smaller model that fits over a larger model that spills to disk.
- Evaluate on your own prompts; public scores rarely predict your workload.
- Log tool calls and retrieval results, not just final answers.
From first model to production and hybrid
Start small: one runtime, one quantized model, one real task. Prove quality and latency before adding retrieval, tools or multi-user serving. Then harden the setup with authentication, network restrictions, sandboxed tool execution, encrypted storage and request logging. Add retrieval only when the task needs external knowledge, and add agents only when a single model call cannot complete the job.
Production local AI usually becomes hybrid. Local inference handles private, routine or high-volume work; a hosted API absorbs hard reasoning, spikes and workloads that exceed local memory. Plugsky serves 30+ models behind an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live, plus cloud, VPC, on-prem and air-gapped deployment options. Audio, image, moderation, batch and fine-tuning endpoints are coming soon. Review plans on the live pricing page.
Honest comparison
| Approach | Setup cost | Privacy | Elasticity | Best for |
|---|---|---|---|---|
| Fully local | Hardware purchase | Highest | Limited to your machine | Private and offline workloads |
| Hybrid local plus cloud | Hardware plus plan | Policy-controlled | Elastic | Mixed sensitivity and load |
| Fully hosted API | Subscription | Provider-processed | Elastic | Fast starts and spiky traffic |
| CPU-only local | Lowest | High | Very limited | Small models and batch jobs |
Frequently asked questions
What is local AI in simple terms?
It means the model runs on hardware you control rather than a provider's servers. Prompts and outputs stay on your machine, and you own the runtime, model and updates.
What hardware do I need to start?
An 8 GB GPU or an Apple Silicon machine with 16 GB or more unified memory runs small and mid-size quantized models. CPU-only works for small models and batch use.
Which runtime should I choose first?
Ollama or LM Studio for the quickest start. Move to llama.cpp for quantization and offload control, or vLLM for concurrent GPU serving.
How much does local AI cost?
You pay for hardware, electricity and operations rather than per token. Compare against hosted plans on the live pricing page once you know your real usage pattern.
Is local AI as good as cloud models?
Small local models are strong at routine tasks; frontier hosted models still lead on hard reasoning and very long context. Hybrid routing gets both.
Can local AI work fully offline?
Yes, if the model, tools and data are local. Features that depend on external services, such as web search, will not work without connectivity.
How do I secure a local AI setup?
Authenticate access, avoid exposing inference ports to the internet, sandbox tool execution, encrypt storage and keep logs free of sensitive text where possible.
When should I stop running locally?
When concurrency, context length or model size exceeds your hardware, or when operating costs exceed a hosted plan. An OpenAI-compatible API keeps the move cheap.