Key facts
| Runtimes | Ollama, LM Studio, llama.cpp and MLX all run models locally |
| Acceleration | Metal GPU backend accelerates llama.cpp and MLX on Apple Silicon |
| Memory model | Unified memory is shared by the CPU, GPU and the operating system |
| Model formats | GGUF for llama.cpp/Ollama; MLX and Safetensors for MLX and LM Studio |
| Practical sizing | 3B-8B models at 4-bit are the comfortable range on 8-16 GB machines |
| Context cost | KV cache grows with context length and competes with weights for memory |
| Hybrid fallback | Plugsky serves 30+ models over one OpenAI-compatible API |
| Endpoint status | Chat, streaming, JSON mode, function calling and embeddings are live |
TL;DR
- Use Ollama or LM Studio for the fastest start; MLX for Apple-native work.
- Unified memory sets the ceiling, so plan weights plus KV cache, not weights alone.
- 4-bit GGUF models are the default trade-off on 8-16 GB Macs.
- Keep prompts and data on-device, then add a cloud fallback for hard tasks.
- An OpenAI-compatible local server keeps code portable to a hosted API later.
How it works, step by step
- Check total unified memory and subtract a few GB for macOS and your apps.
- Install a runtime: Ollama, LM Studio, llama.cpp or MLX.
- Pull a 4-bit quantized model that fits the budget (3B-8B on most laptops).
- Raise context length gradually and watch memory pressure in Activity Monitor.
- Benchmark your real prompts, not generic ones, and note failure cases.
- Expose the runtime on its OpenAI-compatible endpoint so apps share one interface.
- Add a hosted OpenAI-compatible endpoint as fallback for tasks the Mac cannot handle.
Try it yourself
Open the LLM VRAM calculator →
How local inference works on Apple Silicon
Apple Silicon uses unified memory, so the CPU, GPU and Neural Engine draw from the same pool instead of separate VRAM. llama.cpp reaches the GPU through Metal, and MLX is Apple's own array framework built on the same idea. That is why a MacBook can run a capable model without a discrete graphics card.
The consequence is budgeting. Every gigabyte assigned to weights, KV cache, embeddings and the runtime comes out of one number. When memory pressure rises, macOS swaps and generation slows sharply, so a model that fits with headroom beats a larger model that barely fits.
Choosing a model and quantization for your Mac
Start with the smallest model that passes your task, not the largest that fits. A 3B-4B model at 4-bit handles summarization, extraction and simple chat on 8 GB; 7B-8B at 4-bit is the sweet spot on 16 GB; 13B-30B models need 24 GB or more and often a mix of GPU and CPU offload.
- Q4_K_M GGUF files are the usual default: roughly a quarter of fp16 size with modest quality loss.
- Q5 and Q6 are worth it when memory allows and your evaluation shows a real gap.
- Q8 is near-lossless but removes most of the memory savings.
Long context multiplies the KV cache, so cap context when you do not need it and consider KV cache quantization on runtimes that support it.
When to go hybrid with a private API
Macs are excellent for private, offline and bursty work, but sustained throughput and very long contexts hit unified-memory limits. The practical pattern is hybrid: keep sensitive or offline steps local, and route heavy reasoning to an OpenAI-compatible cloud endpoint from the same code path.
Plugsky's API is OpenAI-compatible, so switching between a local server and the hosted endpoint is a base URL change. Chat, streaming, JSON mode, function calling and embeddings are live; audio, image, batch and fine-tuning endpoints are coming soon. See the pricing page for plan details, and start on the free plan with plugsky-micro and plugsky-lite.
Honest comparison
| Concern | Local on macOS | Plugsky private endpoint | Check before deciding |
|---|---|---|---|
| Data path | Prompts stay on the device | Requests go to your chosen deployment | Whether data may leave the machine |
| Model ceiling | Limited by unified memory | 30+ models on one API | Largest model your tasks need |
| Offline use | Works with no network at all | Needs connectivity or a private link | How often you are offline |
| Throughput | One machine, bursty capacity | Scales with the service | Concurrency at peak |
| Operations | You install, patch and monitor | Managed, with SLA and status page | Team time available |
Frequently asked questions
Can a Mac run a local LLM without a discrete GPU?
Yes. Apple Silicon GPUs are used through Metal, and llama.cpp or MLX can also fall back to CPU. Unified memory is the main constraint, not the absence of a graphics card.
How much memory do I need for a 7B model?
A 7B-8B model at 4-bit needs roughly 4-5 GB for weights plus KV cache, runtime overhead and system headroom, so 8 GB is tight and 16 GB is comfortable.
Ollama or LM Studio on macOS?
Ollama is a CLI and background service with a REST API; LM Studio is a desktop app with a GUI and its own local server. Both wrap llama.cpp and expose OpenAI-compatible endpoints.
Does MLX beat llama.cpp?
They are close enough that model and quantization choice matter more than the engine. MLX is Apple-native and evolving quickly; llama.cpp is more portable and supports a wider range of models.
Can I keep using my Mac while a model is loaded?
Small models leave room for normal work, but a large model plus long context can pressure memory and slow the whole system. Watch memory pressure in Activity Monitor.
How do I move from a local Mac setup to a hosted API?
Point your client at an OpenAI-compatible base URL and map model names. The request shape stays the same, so you can keep the local runtime as an offline fallback.
Is local AI on macOS private by default?
Inference is local, but the runtime, your app and any plug-ins can still make network calls. Audit egress and disable telemetry if privacy is the goal.