Local AI

What is the difference between GPTQ and AWQ?

GPTQ and AWQ are both 4-bit post-training quantization methods for GPU inference. GPTQ quantizes layer by layer, compensating for rounding error using calibration data. AWQ scales the weight channels that matter most based on activation statistics before rounding, which protects salient features. Both are supported by modern GPU serving stacks; quality is close, and the right choice depends on the checkpoint and runtime you can use.

Key facts

Bit widthBoth target 4-bit weights, about a quarter of FP16 memory
GPTQ methodLayer-wise quantization with error compensation using a calibration set
AWQ methodActivation-aware per-channel scaling to protect salient weights
HardwareBoth require GPU inference and optimized kernels
Runtime supportvLLM and other serving stacks load GPTQ and AWQ checkpoints
CalibrationQuality depends on the data and implementation used to produce the checkpoint
Cloud optionPlugsky serves 30+ models, so quantization stays behind the API

TL;DR

  • Both methods reach 4-bit with small quality loss on typical workloads.
  • AWQ protects salient channels; GPTQ compensates error layer by layer.
  • Prefer whichever format your serving runtime and target model both support.
  • Checkpoint quality depends on calibration data, not just the method name.
  • Test structured output and long context, where quantization damage appears first.

How it works, step by step

  1. Confirm your serving stack supports the method you are considering.
  2. Find GPTQ and AWQ checkpoints for the exact model version you want.
  3. Record an FP16 or 8-bit baseline on a fixed evaluation set.
  4. Compare the two 4-bit checkpoints on quality, memory and tokens per second.
  5. Test function calling, JSON mode and long-context reading explicitly.
  6. Pick the checkpoint that holds your quality bar at the lowest memory cost.
  7. Keep the evaluation set and source model for future re-quantization.
1Confirm yourserving stacksupports the method2Find GPTQ and AWQcheckpoints for theexact model version3Record an FP16 or8-bit baseline on afixed evaluation4Compare the two4-bit checkpointson quality, memory5Test functioncalling, JSON modeand long-context6Pick the checkpointthat holds yourquality bar at the

Try it yourself

Open the quantization calculator →

Two ways to survive 4-bit rounding

Quantization error is not uniform. Some weights matter far more than others, and naive rounding hurts the ones that carry the most signal. GPTQ addresses this by quantizing one layer at a time and using a calibration set to estimate the error each rounding decision introduces, then adjusting remaining weights to compensate.

AWQ takes a different route. It observes activations and identifies channels whose activation magnitudes are large, then scales those channels up before quantization and scales the corresponding activations down afterward. The scaling makes important weights easier to represent without losing range. Both methods keep most of the model's behaviour at 4-bit, measured by perplexity and task accuracy, though results vary by model and data.

Where they differ in practice

In serving, the differences are mostly tooling. Both formats have GPU kernels in common inference servers, and both support batched requests. Published comparisons often find AWQ slightly ahead at 4-bit on instruction-following quality and GPTQ ahead in older or more custom checkpoint availability, but the margin is small and varies by model.

  • Checkpoint availability: a good checkpoint of one method beats a bad checkpoint of the other.
  • Calibration data: a checkpoint quantized on data close to your domain tends to serve your domain better.
  • Runtime kernels: verify support for the exact model architecture, not just the method.

If your stack supports both and quality is close, benchmark throughput at your real batch size and context length. Sequence length affects the KV cache more than the weight format, so include it in the test.

How to choose and when to skip the choice

Choose by constraint order: runtime support first, then checkpoint availability for the exact model version, then quality on your evaluation set. Do not choose on method reputation alone, because the implementation and calibration matter as much as the technique.

Move up in precision when quality regresses: an 8-bit model of the same family may beat a 4-bit model you would need to change. Move down when memory forces the issue, and accept that a larger model at 4-bit often outperforms a smaller model at 8-bit for the same memory budget.

If you would rather not manage checkpoints and kernels, a hosted API removes the entire decision. Plugsky serves 30+ models over an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. See the live pricing page for current plans.

Honest comparison

FactorGPTQAWQ8-bit alternative
MethodLayer-wise error compensationActivation-aware scalingHigher-precision quantization
Bit widthTypically 4-bitTypically 4-bit8-bit
Memory vs FP16About a quarterAbout a quarterAbout half
Quality riskDepends on calibrationDepends on calibrationLower at higher memory
Runtime supportWide GPU stack supportWide GPU stack supportWidest

Frequently asked questions

Which is more accurate, GPTQ or AWQ?

At 4-bit they are close, and results vary by model and calibration data. Test both checkpoints of the same model on your own evaluation set rather than relying on general rankings.

Can AWQ models run on any GPU?

They need GPU kernels supported by your runtime and hardware. Check the serving stack's compatibility list for the exact architecture.

Do both support function calling?

The method does not determine tool-calling support; the model and runtime do. Quantization can increase malformed structured output, so verify after quantizing.

Should I choose 4-bit or 8-bit?

Use 4-bit when memory is tight and the larger model matters; use 8-bit when you have headroom and want a smaller quality gap.

Does calibration data matter?

Yes. Checkpoints quantized on data closer to your workload often perform better on that workload, especially for specialized domains.

Can I quantize with GPTQ or AWQ myself?

Yes, both have open tooling pipelines, but you need calibration data, GPU time and a way to validate the result. Using a well-made public checkpoint is often faster.

Which format does vLLM support?

vLLM supports multiple 4-bit formats including GPTQ and AWQ, but support is version and architecture dependent. Confirm before planning a deployment.