Key facts
| Q4 size | Roughly half a gigabyte per billion parameters plus overhead |
| Q8 size | Roughly one gigabyte per billion parameters |
| Speed | Lower-bit weights move less data per token, helping bandwidth-bound hardware |
| Quality | Q8 is near-lossless; Q4 shows a small, task-dependent drop |
| Variants | Q4_K_M and Q8_0 are common GGUF presets with different mixes |
| KV cache | Cache precision is a separate setting from weight quantization |
| When Q8 fits | Smaller models on roomy GPUs benefit most from the higher level |
| Hosted option | Plugsky serves 30+ hosted models over one OpenAI-compatible API |
TL;DR
- Q8 uses about twice the memory of Q4 for a small quality gain.
- Q4 is the sensible default when memory is the constraint.
- Q8 is worth it when the model fits comfortably and fidelity matters.
- Weight quantization and KV cache precision are separate settings.
- Measure on your evaluation set; the gap depends on the task.
How it works, step by step
- Compute the memory budget for your hardware after system overhead.
- Run your model at Q4_K_M and record quality on a fixed evaluation set.
- Run the same model at Q8_0 and record the same metrics.
- Compare memory use, speed and failure cases side by side.
- Choose the lowest bit width that passes your quality bar.
- Tune KV cache precision separately if context is the bottleneck.
- Pin the exact file and revision you validated.
Original data
Try it yourself
Open the quantization calculator →
What Q4 and Q8 actually mean
Quantization reduces the precision used to store model weights. Q4 keeps roughly four bits per weight and Q8 roughly eight, which is why memory use differs by about a factor of two. The numbers in a preset name such as Q4_K_M describe how different tensors are treated, not one uniform bit width.
Lower precision shrinks the model and reduces the bytes moved per generated token. That is why quantization often speeds up generation on memory-bandwidth-limited hardware, while on compute-heavy setups the difference is smaller.
Quality, memory and speed in practice
Q8 is close enough to full precision that most teams treat it as near-lossless. Q4 introduces a small quality drop that varies by model and task; summarization and chat usually hold up well, while tasks requiring exact formatting, arithmetic or careful instruction following show more sensitivity.
- Memory: Q4 fits models that Q8 cannot, which often decides the choice outright.
- Speed: Q4 usually generates faster where bandwidth is the limit.
- Quality: Q8 is safer for high-fidelity work when the model fits.
- KV cache: cache precision is separate, so long context can be tuned independently.
How to choose and verify
Start from the memory budget. If the model fits at Q8 with context headroom, use Q8. If it does not, use Q4_K_M and check whether it passes your quality bar. Only move to lower bit widths when memory forces it, because quality loss grows and variant behaviour varies more.
Then verify with your own evaluation set rather than generic impressions. Keep a fixed prompt list, score task success and format validity, and pin the exact model file you validated so deployments are reproducible. If local memory remains the constraint, a hosted OpenAI-compatible API serves 30+ models without any quantization decisions on your side. See pricing for plans and start free with plugsky-micro and plugsky-lite.
Honest comparison
| Concern | Q4 | Q8 | Check before deciding |
|---|---|---|---|
| Memory | Roughly half of Q8 | Roughly double Q4 | Model size versus available memory |
| Quality | Small task-dependent drop | Near-lossless | Your evaluation set |
| Speed | Usually faster where bandwidth-bound | Slower per token | Hardware profile |
| Use when | Memory or speed is tight | The model fits with headroom | Budget |
| Typical preset | Q4_K_M | Q8_0 | Runtime support |
Frequently asked questions
Is Q4 much worse than Q8?
On most everyday tasks the difference is small and often invisible, but it depends on the model and workload. Tasks needing precise formatting or exact recall are more sensitive, so test before deciding.
How much memory does each level need?
As a rule, 4-bit weights are roughly half a gigabyte per billion parameters and 8-bit roughly one gigabyte, before KV cache and runtime overhead.
Why is Q4 sometimes faster?
Generation is often limited by memory bandwidth, so smaller weights mean less data moved per token. On compute-bound hardware the gap narrows.
Should I use an even lower bit width?
Below four bits, quality loss becomes more likely and variant behaviour varies more. Use it only when memory forces it and validate carefully.
Does quantization affect the KV cache?
No, cache precision is a separate setting. You can quantize weights and leave the cache at higher precision, or the reverse, depending on what is tight.
Do hosted models use quantization?
Provider-side serving may use optimized formats internally, but you do not manage them. A hosted endpoint sidesteps local memory limits entirely.
What is the safest default?
Start at Q4_K_M. If quality fails on your evaluation set and the model fits at Q8, move up; if it fits at Q8 from the start, use it.