Key facts
| Bit width | Both target 4-bit weights, about a quarter of FP16 memory |
| GPTQ method | Layer-wise quantization with error compensation using a calibration set |
| AWQ method | Activation-aware per-channel scaling to protect salient weights |
| Hardware | Both require GPU inference and optimized kernels |
| Runtime support | vLLM and other serving stacks load GPTQ and AWQ checkpoints |
| Calibration | Quality depends on the data and implementation used to produce the checkpoint |
| Cloud option | Plugsky serves 30+ models, so quantization stays behind the API |
TL;DR
- Both methods reach 4-bit with small quality loss on typical workloads.
- AWQ protects salient channels; GPTQ compensates error layer by layer.
- Prefer whichever format your serving runtime and target model both support.
- Checkpoint quality depends on calibration data, not just the method name.
- Test structured output and long context, where quantization damage appears first.
How it works, step by step
- Confirm your serving stack supports the method you are considering.
- Find GPTQ and AWQ checkpoints for the exact model version you want.
- Record an FP16 or 8-bit baseline on a fixed evaluation set.
- Compare the two 4-bit checkpoints on quality, memory and tokens per second.
- Test function calling, JSON mode and long-context reading explicitly.
- Pick the checkpoint that holds your quality bar at the lowest memory cost.
- Keep the evaluation set and source model for future re-quantization.
Try it yourself
Open the quantization calculator →
Two ways to survive 4-bit rounding
Quantization error is not uniform. Some weights matter far more than others, and naive rounding hurts the ones that carry the most signal. GPTQ addresses this by quantizing one layer at a time and using a calibration set to estimate the error each rounding decision introduces, then adjusting remaining weights to compensate.
AWQ takes a different route. It observes activations and identifies channels whose activation magnitudes are large, then scales those channels up before quantization and scales the corresponding activations down afterward. The scaling makes important weights easier to represent without losing range. Both methods keep most of the model's behaviour at 4-bit, measured by perplexity and task accuracy, though results vary by model and data.
Where they differ in practice
In serving, the differences are mostly tooling. Both formats have GPU kernels in common inference servers, and both support batched requests. Published comparisons often find AWQ slightly ahead at 4-bit on instruction-following quality and GPTQ ahead in older or more custom checkpoint availability, but the margin is small and varies by model.
- Checkpoint availability: a good checkpoint of one method beats a bad checkpoint of the other.
- Calibration data: a checkpoint quantized on data close to your domain tends to serve your domain better.
- Runtime kernels: verify support for the exact model architecture, not just the method.
If your stack supports both and quality is close, benchmark throughput at your real batch size and context length. Sequence length affects the KV cache more than the weight format, so include it in the test.
How to choose and when to skip the choice
Choose by constraint order: runtime support first, then checkpoint availability for the exact model version, then quality on your evaluation set. Do not choose on method reputation alone, because the implementation and calibration matter as much as the technique.
Move up in precision when quality regresses: an 8-bit model of the same family may beat a 4-bit model you would need to change. Move down when memory forces the issue, and accept that a larger model at 4-bit often outperforms a smaller model at 8-bit for the same memory budget.
If you would rather not manage checkpoints and kernels, a hosted API removes the entire decision. Plugsky serves 30+ models over an OpenAI-compatible endpoint with chat, streaming, JSON mode, function calling, embeddings, RAG and agents live; audio, image, moderation, batch and fine-tuning endpoints are coming soon. See the live pricing page for current plans.
Honest comparison
| Factor | GPTQ | AWQ | 8-bit alternative |
|---|---|---|---|
| Method | Layer-wise error compensation | Activation-aware scaling | Higher-precision quantization |
| Bit width | Typically 4-bit | Typically 4-bit | 8-bit |
| Memory vs FP16 | About a quarter | About a quarter | About half |
| Quality risk | Depends on calibration | Depends on calibration | Lower at higher memory |
| Runtime support | Wide GPU stack support | Wide GPU stack support | Widest |
Frequently asked questions
Which is more accurate, GPTQ or AWQ?
At 4-bit they are close, and results vary by model and calibration data. Test both checkpoints of the same model on your own evaluation set rather than relying on general rankings.
Can AWQ models run on any GPU?
They need GPU kernels supported by your runtime and hardware. Check the serving stack's compatibility list for the exact architecture.
Do both support function calling?
The method does not determine tool-calling support; the model and runtime do. Quantization can increase malformed structured output, so verify after quantizing.
Should I choose 4-bit or 8-bit?
Use 4-bit when memory is tight and the larger model matters; use 8-bit when you have headroom and want a smaller quality gap.
Does calibration data matter?
Yes. Checkpoints quantized on data closer to your workload often perform better on that workload, especially for specialized domains.
Can I quantize with GPTQ or AWQ myself?
Yes, both have open tooling pipelines, but you need calibration data, GPU time and a way to validate the result. Using a well-made public checkpoint is often faster.
Which format does vLLM support?
vLLM supports multiple 4-bit formats including GPTQ and AWQ, but support is version and architecture dependent. Confirm before planning a deployment.