Local AI

What is the difference between GGUF and GPTQ?

GGUF is llama.cpp's portable model container that runs on CPU, Metal and GPUs with many quantization levels. GPTQ is a GPU-oriented post-training quantization method that uses calibration data to minimize rounding error, typically at 4-bit. Choose GGUF for broad hardware support and tunable quality; choose GPTQ when you serve on GPUs with runtimes that load GPTQ checkpoints.

Key facts

Format typeGGUF is a container; GPTQ is a quantization method with GPU checkpoint layouts
HardwareGGUF runs on CPU, Metal, CUDA, ROCm and Vulkan; GPTQ requires a GPU
MethodGPTQ quantizes layer by layer using Hessian-based error compensation on calibration data
Typical bitsGPTQ is commonly 4-bit; GGUF spans roughly 2-bit to 8-bit levels
Runtimesllama.cpp loads GGUF; vLLM and ExLlama-style stacks load GPTQ
Memory at 4-bitBoth reduce weight memory to about a quarter of FP16
Cloud optionPlugsky serves 30+ models over an OpenAI-compatible API

TL;DR

  • GGUF is the compatibility choice across CPU, Apple Silicon and GPUs.
  • GPTQ is the GPU-native choice for batched serving stacks.
  • GPTQ quality depends on the calibration dataset used at quantization time.
  • GGUF lets you step quality up or down with different quantization levels.
  • Both formats cut 4-bit weight memory to roughly a quarter of FP16.

How it works, step by step

  1. List the hardware and runtime that will serve the model.
  2. For CPU or Apple Silicon, choose a GGUF build and a Q4 or Q5 level.
  3. For GPU serving, check whether your runtime prefers GPTQ or another GPU format.
  4. Establish an FP16 baseline on a fixed evaluation set.
  5. Compare quantized quality, memory and throughput at your target context length.
  6. Verify structured output such as JSON and tool calls after quantizing.
  7. Store the evaluation set so you can re-quantize when models or hardware change.
1List the hardwareand runtime thatwill serve the2For CPU or AppleSilicon, choose aGGUF build and a Q43For GPU serving,check whether yourruntime prefers4Establish an FP16baseline on a fixedevaluation set.5Compare quantizedquality, memory andthroughput at your6Verify structuredoutput such as JSONand tool calls

Original data

GPTQ is commonTypical bitsBoth reduce weMemory at 4-bitPlugsky servesCloud optionSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the model quantization selector →

How the two approaches differ

GGUF describes a file: weights, tokenizer and metadata in one memory-mappable container that llama.cpp can run almost anywhere. Its quantization levels range from very aggressive 2-3 bit variants to near-lossless 8-bit options, and the K-quant families allocate bits unevenly to protect sensitive tensors.

GPTQ describes how weights were quantized. It processes layers in order, using a small calibration set to estimate the error introduced by rounding and compensate for it. The result is usually a 4-bit GPU checkpoint. Because calibration data shapes the result, quality varies by who produced the checkpoint and with what data.

Hardware and runtime fit

The deployment target decides. GGUF is the only realistic option for CPU-only machines and the most convenient choice on Apple Silicon via Metal. It also works on GPUs with layer offload, which is useful when a model only partly fits in VRAM. GPTQ requires a supported GPU and a runtime with GPTQ kernels; it does not run on CPU.

  • Single desktop or laptop: GGUF, with offload as needed.
  • Apple Silicon: GGUF through Metal, or MLX-native formats.
  • GPU server with batching: GPTQ or another GPU-native format supported by your stack.
  • Mixed fleet: GGUF for edge and CPU nodes, GPU formats for servers.

Check runtime support for the exact model architecture, since kernel availability lags new architectures.

Quality and practical selection

Well-made 4-bit quantizations of either format are close to FP16 on most tasks, and the gap widens below 4-bit. GGUF gives you a ladder: if Q4_K_M shows problems, try Q5_K_M or Q6_K before changing models. GPTQ gives you fewer knobs per checkpoint, so the producer's calibration quality matters more.

Test structured behaviour explicitly. Function calling, JSON mode and long-context reading are where quantization damage tends to surface first, and those capabilities matter more than perplexity for real workloads. Fix a prompt set, record results, and re-run it whenever you change format, level or runtime version. If managing the stack is not the point of your project, Plugsky removes the choice: 30+ models behind an OpenAI-compatible API, with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live, while audio, image, moderation, batch and fine-tuning endpoints are coming soon. Plan details are on the live pricing page.

Honest comparison

FactorGGUFGPTQFP16 checkpoint
TypeContainer formatQuantization methodFull precision
HardwareCPU, Metal, CUDA, ROCm, VulkanGPU onlyGPU
Bit widthsAbout 2-bit to 8-bitTypically 4-bit16-bit
Runtimellama.cpp and derivativesvLLM, ExLlama-style stacksMost runtimes
Memory at 4-bitAbout a quarter of FP16About a quarter of FP16Reference

Frequently asked questions

Can GPTQ run on CPU?

No. GPTQ checkpoints need GPU kernels. Use GGUF if CPU inference is required.

Is GPTQ more accurate than GGUF at 4-bit?

They are comparable when both are well produced. GPTQ quality depends on calibration data and the implementation used to create the checkpoint.

Which GGUF level matches GPTQ 4-bit?

Q4_K_M is the closest common equivalent. Q5_K_M and Q6_K trade more memory for better fidelity.

Do both formats support the same model architectures?

Support varies. Popular architectures are usually available in both, but new releases may appear in one format first.

Does quantization break function calling?

It can increase malformed structured output at aggressive bit widths. Test tool calling and JSON mode on the exact checkpoint you plan to deploy.

Which format should I use for a GPU server with many users?

A GPU-native format supported by your batching runtime, commonly GPTQ or AWQ, is the usual choice for concurrent traffic.

Do I need to care about formats when using a hosted API?

No. Plugsky runs inference for you across 30+ models through an OpenAI-compatible endpoint, so quantization is an internal detail.