Local AI

What hardware do you need to run AI locally?

Local AI hardware is chosen by memory first, bandwidth second and compute third. GPUs with 8-24 GB of VRAM cover consumer local models; Apple Silicon offers large unified memory for its price class; CPU rigs run small quantized models without a GPU. The practical rule is that model weights plus KV cache must fit in fast memory, and bandwidth sets how fast tokens appear.

Key facts

Primary constraintFast memory capacity decides which models fit
Secondary constraintMemory bandwidth sets tokens per second
Consumer GPUs8 GB, 12 GB, 16 GB and 24 GB VRAM classes are common entry points
Apple SiliconUnified memory is shared across CPU and GPU, so larger models can fit
CPU onlyllama.cpp runs quantized models with no GPU, at lower speed
StorageNVMe matters because models are tens of gigabytes and load from disk
Cloud optionPlugsky serves 30+ models so you can avoid buying hardware for every workload

TL;DR

  • Fit matters more than raw compute: if weights and KV cache do not fit, speed is irrelevant.
  • 8 GB runs small models; 16 GB is a comfortable middle; 24 GB handles 30B-class models at 4-bit.
  • Apple Silicon trades peak speed for large unified memory and low power draw.
  • A CPU-only machine can serve small quantized models for batch or occasional use.
  • Buy for your baseline workload and rent or route overflow to a hosted API.

How it works, step by step

  1. Estimate model memory: parameters times bits per weight, then add KV cache and overhead.
  2. Decide which model sizes you need, not which model you wish you could run.
  3. Compare hardware classes by memory capacity, bandwidth and power draw.
  4. Check software support: your runtime must support the GPU vendor and OS.
  5. Plan storage for several model files and a vector index.
  6. Benchmark your real prompts and context length before committing to a purchase.
  7. Reserve a cloud or hosted API route for workloads that exceed local capacity.
1Estimate modelmemory: parameterstimes bits per2Decide which modelsizes you need, notwhich model you3Compare hardwareclasses by memorycapacity, bandwidth4Check softwaresupport: yourruntime must5Plan storage forseveral model filesand a vector index.6Benchmark your realprompts and contextlength before

Try it yourself

Open the GPU fit checker →

The memory rule and why bandwidth matters

Every local model has a memory floor: weights plus KV cache plus runtime overhead. At 4-bit, a 7B model needs roughly 4-5 GB for weights, a 14B model around 8-9 GB, and a 32B model around 17-19 GB. Context adds to that through the KV cache, so a model that fits at 4k tokens may not fit at 32k.

Once it fits, speed is mostly bandwidth. Generating a token requires reading weights from memory, so faster memory produces more tokens per second. That is why a GPU with moderate compute but high bandwidth can beat a stronger compute part with narrow memory. Capacity gets you running; bandwidth gets you usable.

Hardware classes compared

Consumer GPUs are the common choice. The 8 GB class handles small models and tight 7B-8B quantization; 12-16 GB runs 7B-14B models comfortably; 24 GB handles 30B-class models at low-bit quantization and 14B models at higher precision. VRAM is dedicated and bandwidth is high, but capacity caps the model size.

Apple Silicon shares one memory pool between CPU and GPU, so a 32 GB or 64 GB machine can hold models that would need an expensive GPU. Bandwidth is lower than discrete high-end GPUs, so prompt processing and generation are slower, but power draw and noise are far lower, which suits laptops and quiet offices.

CPU-only rigs are the budget path. Modern CPUs with wide vector instructions run quantized models through llama.cpp, but throughput suits batch jobs, summarisation pipelines and personal use rather than interactive multi-user traffic. Maximise memory bandwidth with dual-channel or multi-channel configurations.

Buying strategy and hybrid routing

Buy for the workload you run every day and route the rest. A single 16 GB or 24 GB GPU covers most personal and small-team local inference; specialist workloads such as 70B-class models, heavy fine-tuning or high concurrency are usually cheaper to run in the cloud than to host on owned hardware, once power and operations are counted.

Keep the interface portable so routing is easy. An OpenAI-compatible endpoint means the same application code can run against a local server or a hosted API. Plugsky hosts 30+ models behind that interface with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. Compare cost before buying: the live pricing page covers hosted plans, and calculators help you test assumptions.

Honest comparison

ClassTypical memoryModel range at 4-bitStrengthsTrade-offs
Entry GPU8 GB3B-8BLow cost, good speedTight context, small models
Mid GPU12-16 GB7B-14BBalanced for daily useUpper model sizes need offload
High-end GPU24 GB14B-32BFast, larger modelsCost, power, capacity ceiling
Apple Silicon16-128 GB unified3B-70B depending on RAMLarge memory, quiet, efficientLower bandwidth than top GPUs
CPU-onlySystem RAMSmall quantized modelsNo GPU purchaseSlow interactive use

Frequently asked questions

How much VRAM do I need for a 7B model?

Roughly 4-5 GB for 4-bit weights, plus KV cache and overhead. 8 GB is the realistic floor; 12 GB or more gives comfortable context.

Is a GPU always faster than Apple Silicon?

For raw generation speed on comparable memory, discrete high-bandwidth GPUs usually lead. Apple Silicon compensates with large unified memory and much lower power draw.

Can I run AI with no GPU at all?

Yes. llama.cpp runs quantized models on CPU. It is viable for small models, batch work and occasional queries, but not for high-concurrency interactive use.

How much system RAM do I need?

Match RAM to the models you run. CPU inference needs the full model plus cache in system memory, so 16 GB is a floor for small models and 32-64 GB is common for larger ones.

Does SSD speed matter?

Loading multi-gigabyte model files from disk is faster on NVMe, and retrieval indexes benefit too. Once loaded, generation speed depends on memory bandwidth.

Should I buy hardware or use a hosted API?

If your workload is steady and fits one GPU, buying can pay off. For spiky, high-concurrency or frontier-model workloads, hosted or hybrid routing is usually simpler and cheaper to operate.

Can I mix a local machine with a cloud API?

Yes, and it is common. Keep local inference for private and routine work, and route heavier traffic to an OpenAI-compatible hosted API without changing application code.

What about multi-GPU setups?

Multiple GPUs can pool memory for larger models, but interconnect and software support determine how well they scale. Check runtime support and real benchmarks before investing.