Key facts
| Memory model | Unified memory shared by CPU and GPU; capacity varies by configuration |
| Runtimes | llama.cpp with Metal, MLX, Ollama and LM Studio |
| Model format | GGUF through llama.cpp or Ollama; MLX-native models for MLX |
| Speed driver | Memory bandwidth, which scales with chip tier |
| Model sizing | 7B at 4-bit needs roughly 4-5 GB; capacity is set by installed RAM |
| Power | Low power draw and quiet operation compared with GPU workstations |
| Cloud option | Plugsky serves 30+ hosted models over an OpenAI-compatible API |
TL;DR
- Unified memory is the headline: a 32GB or 64GB Mac holds models that need an expensive discrete GPU.
- Bandwidth, not compute, determines tokens per second, so chip tier matters.
- llama.cpp with Metal is the compatibility default; MLX is the Apple-native fast path.
- Leave 8-16 GB of RAM for the system when sizing a model.
- Use a hosted API for frontier models or heavy concurrency.
How it works, step by step
- Check installed memory and reserve 8-16 GB for macOS and applications.
- Size the model: 7B at 4-bit needs roughly 4-5 GB of weights plus cache.
- Install a runtime: Ollama or LM Studio for ease, llama.cpp or MLX for control.
- Download a quantized GGUF or an MLX model that fits your budget.
- Serve the OpenAI-compatible endpoint and test with your normal prompts.
- Measure tokens per second at your real context length, not a toy prompt.
- Add a hosted fallback for models too large for the machine.
Try it yourself
Open the local model recommender →
Why unified memory changes the equation
On a discrete GPU, model weights must fit in dedicated VRAM, and system RAM does not help. Apple Silicon shares one memory pool between CPU and GPU, so a machine with more RAM can hold larger models than a GPU with a fixed memory ceiling. That is why Macs became popular for local inference: a mid-range laptop can run models that would require a high-end graphics card.
The trade-off is bandwidth. Token generation is memory-bound, so speed scales with the chip tier and its memory bandwidth rather than with core count alone. A model may fit comfortably and still generate slowly if the machine is bandwidth-limited. Fit is about capacity; speed is about bandwidth.
Runtimes: llama.cpp, MLX and the rest
llama.cpp is the compatibility choice. Its Metal backend accelerates GGUF models on Apple Silicon, and the same format works on Linux and Windows machines, which keeps workflows portable. Ollama wraps a similar stack with simple model management, and LM Studio provides a GUI plus a local server.
MLX is Apple's framework for machine learning on Apple silicon and has become a fast path for many model families, with its own model conversions. It often performs well on Apple hardware and integrates neatly with Python workflows, at the cost of a narrower model ecosystem than GGUF. Many users keep both: GGUF for breadth and MLX for the models that run best there.
- Ease: Ollama or LM Studio.
- Portability: GGUF via llama.cpp.
- Apple-native performance: MLX.
Sizing models to memory and staying hybrid
Start by reserving memory for the system. macOS and your applications need several gigabytes, and heavy swap destroys generation speed. With 16GB of RAM, small and mid-size models at 4-bit are realistic. With 32GB, 14B-class models at 4-bit fit with context. With 64GB and above, 30B-class models and aggressive quantization of larger ones become possible.
Context still costs cache on top of weights, and long conversations or retrieved passages can push a comfortable setup into swap. Cap context, summarise history and keep retrieval tight. When a task needs a model too large for the machine, route it to a hosted API. Plugsky serves 30+ models behind an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. The free plan includes plugsky-micro and plugsky-lite, and a 14-day full-access trial covers evaluation. See the live pricing page for plans.
Honest comparison
| Runtime | Model format | Acceleration | Best for | Ecosystem |
|---|---|---|---|---|
| llama.cpp | GGUF | Metal | Portability and tuning | Very broad |
| MLX | MLX-native | Metal and Apple frameworks | Apple-native performance | Growing |
| Ollama | GGUF | Metal | Simple local setup | Broad |
| LM Studio | GGUF and others | Metal | Desktop evaluation | Broad |
Frequently asked questions
Do I need a Mac with a Pro or Max chip?
No. Base chips run small and mid-size quantized models. Higher tiers add memory capacity and bandwidth, which let you run larger models faster.
How much RAM should I have for local AI?
16GB runs small models, 32GB comfortably handles 14B-class models at 4-bit, and 64GB and above opens up 30B-class models. Reserve 8-16 GB for the system.
Is MLX faster than llama.cpp on Macs?
It depends on the model and version. MLX is often strong on Apple hardware, while llama.cpp offers broader model support. Benchmark your exact model on both.
Can I run a 70B model on a Mac?
A high-memory configuration can hold 70B models at 4-bit, but generation is slow because it is bandwidth-limited. Memory pressure from other applications also matters.
Does local AI drain the battery?
Generation uses the GPU heavily and reduces battery life, though less than a comparable discrete-GPU laptop. Keep long jobs plugged in.
Which apps support local models on macOS?
Ollama, LM Studio, Jan and similar apps run local models on Apple Silicon, and any OpenAI-compatible client can connect to their local endpoints.
Can I use a Mac as a shared inference server?
It works for small teams with light traffic. For heavier concurrency, a GPU server or a hosted API scales better.
What about privacy?
Local inference keeps prompts and data on the machine. Any hybrid route you configure determines the rest.