Local AI

How do you run local AI on Apple Silicon?

Apple Silicon runs local AI by sharing one pool of unified memory between CPU and GPU, so large models can fit without a discrete GPU. llama.cpp uses Metal for GPU acceleration; MLX is Apple's own array framework with optimized model ports. Memory capacity sets the model size and bandwidth sets the speed, so higher-tier chips generate faster while lower tiers rely on smaller quantized models.

Key facts

Memory modelUnified memory shared by CPU and GPU; capacity varies by configuration
Runtimesllama.cpp with Metal, MLX, Ollama and LM Studio
Model formatGGUF through llama.cpp or Ollama; MLX-native models for MLX
Speed driverMemory bandwidth, which scales with chip tier
Model sizing7B at 4-bit needs roughly 4-5 GB; capacity is set by installed RAM
PowerLow power draw and quiet operation compared with GPU workstations
Cloud optionPlugsky serves 30+ hosted models over an OpenAI-compatible API

TL;DR

  • Unified memory is the headline: a 32GB or 64GB Mac holds models that need an expensive discrete GPU.
  • Bandwidth, not compute, determines tokens per second, so chip tier matters.
  • llama.cpp with Metal is the compatibility default; MLX is the Apple-native fast path.
  • Leave 8-16 GB of RAM for the system when sizing a model.
  • Use a hosted API for frontier models or heavy concurrency.

How it works, step by step

  1. Check installed memory and reserve 8-16 GB for macOS and applications.
  2. Size the model: 7B at 4-bit needs roughly 4-5 GB of weights plus cache.
  3. Install a runtime: Ollama or LM Studio for ease, llama.cpp or MLX for control.
  4. Download a quantized GGUF or an MLX model that fits your budget.
  5. Serve the OpenAI-compatible endpoint and test with your normal prompts.
  6. Measure tokens per second at your real context length, not a toy prompt.
  7. Add a hosted fallback for models too large for the machine.
1Check installedmemory and reserve8-16 GB for macOS2Size the model: 7Bat 4-bit needsroughly 4-5 GB of3Install a runtime:Ollama or LM Studiofor ease, llama.cpp4Download aquantized GGUF oran MLX model that5Serve theOpenAI-compatibleendpoint and test6Measure tokens persecond at your realcontext length, not

Try it yourself

Open the local model recommender →

Why unified memory changes the equation

On a discrete GPU, model weights must fit in dedicated VRAM, and system RAM does not help. Apple Silicon shares one memory pool between CPU and GPU, so a machine with more RAM can hold larger models than a GPU with a fixed memory ceiling. That is why Macs became popular for local inference: a mid-range laptop can run models that would require a high-end graphics card.

The trade-off is bandwidth. Token generation is memory-bound, so speed scales with the chip tier and its memory bandwidth rather than with core count alone. A model may fit comfortably and still generate slowly if the machine is bandwidth-limited. Fit is about capacity; speed is about bandwidth.

Runtimes: llama.cpp, MLX and the rest

llama.cpp is the compatibility choice. Its Metal backend accelerates GGUF models on Apple Silicon, and the same format works on Linux and Windows machines, which keeps workflows portable. Ollama wraps a similar stack with simple model management, and LM Studio provides a GUI plus a local server.

MLX is Apple's framework for machine learning on Apple silicon and has become a fast path for many model families, with its own model conversions. It often performs well on Apple hardware and integrates neatly with Python workflows, at the cost of a narrower model ecosystem than GGUF. Many users keep both: GGUF for breadth and MLX for the models that run best there.

  • Ease: Ollama or LM Studio.
  • Portability: GGUF via llama.cpp.
  • Apple-native performance: MLX.

Sizing models to memory and staying hybrid

Start by reserving memory for the system. macOS and your applications need several gigabytes, and heavy swap destroys generation speed. With 16GB of RAM, small and mid-size models at 4-bit are realistic. With 32GB, 14B-class models at 4-bit fit with context. With 64GB and above, 30B-class models and aggressive quantization of larger ones become possible.

Context still costs cache on top of weights, and long conversations or retrieved passages can push a comfortable setup into swap. Cap context, summarise history and keep retrieval tight. When a task needs a model too large for the machine, route it to a hosted API. Plugsky serves 30+ models behind an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. The free plan includes plugsky-micro and plugsky-lite, and a 14-day full-access trial covers evaluation. See the live pricing page for plans.

Honest comparison

RuntimeModel formatAccelerationBest forEcosystem
llama.cppGGUFMetalPortability and tuningVery broad
MLXMLX-nativeMetal and Apple frameworksApple-native performanceGrowing
OllamaGGUFMetalSimple local setupBroad
LM StudioGGUF and othersMetalDesktop evaluationBroad

Frequently asked questions

Do I need a Mac with a Pro or Max chip?

No. Base chips run small and mid-size quantized models. Higher tiers add memory capacity and bandwidth, which let you run larger models faster.

How much RAM should I have for local AI?

16GB runs small models, 32GB comfortably handles 14B-class models at 4-bit, and 64GB and above opens up 30B-class models. Reserve 8-16 GB for the system.

Is MLX faster than llama.cpp on Macs?

It depends on the model and version. MLX is often strong on Apple hardware, while llama.cpp offers broader model support. Benchmark your exact model on both.

Can I run a 70B model on a Mac?

A high-memory configuration can hold 70B models at 4-bit, but generation is slow because it is bandwidth-limited. Memory pressure from other applications also matters.

Does local AI drain the battery?

Generation uses the GPU heavily and reduces battery life, though less than a comparable discrete-GPU laptop. Keep long jobs plugged in.

Which apps support local models on macOS?

Ollama, LM Studio, Jan and similar apps run local models on Apple Silicon, and any OpenAI-compatible client can connect to their local endpoints.

Can I use a Mac as a shared inference server?

It works for small teams with light traffic. For heavier concurrency, a GPU server or a hosted API scales better.

What about privacy?

Local inference keeps prompts and data on the machine. Any hybrid route you configure determines the rest.