Local AI

Should you use Q4 or Q8 quantization for a local LLM?

Q4 stores weights at roughly four bits and Q8 at roughly eight, so Q8 uses about twice the memory of Q4 for a small quality gain. Choose Q4 when memory or speed is tight and Q8 when the model fits and output fidelity matters. Test both on your own evaluation set before deciding.

Key facts

Q4 sizeRoughly half a gigabyte per billion parameters plus overhead
Q8 sizeRoughly one gigabyte per billion parameters
SpeedLower-bit weights move less data per token, helping bandwidth-bound hardware
QualityQ8 is near-lossless; Q4 shows a small, task-dependent drop
VariantsQ4_K_M and Q8_0 are common GGUF presets with different mixes
KV cacheCache precision is a separate setting from weight quantization
When Q8 fitsSmaller models on roomy GPUs benefit most from the higher level
Hosted optionPlugsky serves 30+ hosted models over one OpenAI-compatible API

TL;DR

  • Q8 uses about twice the memory of Q4 for a small quality gain.
  • Q4 is the sensible default when memory is the constraint.
  • Q8 is worth it when the model fits comfortably and fidelity matters.
  • Weight quantization and KV cache precision are separate settings.
  • Measure on your evaluation set; the gap depends on the task.

How it works, step by step

  1. Compute the memory budget for your hardware after system overhead.
  2. Run your model at Q4_K_M and record quality on a fixed evaluation set.
  3. Run the same model at Q8_0 and record the same metrics.
  4. Compare memory use, speed and failure cases side by side.
  5. Choose the lowest bit width that passes your quality bar.
  6. Tune KV cache precision separately if context is the bottleneck.
  7. Pin the exact file and revision you validated.
1Compute the memorybudget for yourhardware after2Run your model atQ4_K_M and recordquality on a fixed3Run the same modelat Q8_0 and recordthe same metrics.4Compare memory use,speed and failurecases side by side.5Choose the lowestbit width thatpasses your quality6Tune KV cacheprecisionseparately if

Original data

Q8 is near-losQualityQ4_K_M and Q8_VariantsPlugsky servesHosted optionSource: Plugsky facts table · updated 2026-09-26

Try it yourself

Open the quantization calculator →

What Q4 and Q8 actually mean

Quantization reduces the precision used to store model weights. Q4 keeps roughly four bits per weight and Q8 roughly eight, which is why memory use differs by about a factor of two. The numbers in a preset name such as Q4_K_M describe how different tensors are treated, not one uniform bit width.

Lower precision shrinks the model and reduces the bytes moved per generated token. That is why quantization often speeds up generation on memory-bandwidth-limited hardware, while on compute-heavy setups the difference is smaller.

Quality, memory and speed in practice

Q8 is close enough to full precision that most teams treat it as near-lossless. Q4 introduces a small quality drop that varies by model and task; summarization and chat usually hold up well, while tasks requiring exact formatting, arithmetic or careful instruction following show more sensitivity.

  • Memory: Q4 fits models that Q8 cannot, which often decides the choice outright.
  • Speed: Q4 usually generates faster where bandwidth is the limit.
  • Quality: Q8 is safer for high-fidelity work when the model fits.
  • KV cache: cache precision is separate, so long context can be tuned independently.

How to choose and verify

Start from the memory budget. If the model fits at Q8 with context headroom, use Q8. If it does not, use Q4_K_M and check whether it passes your quality bar. Only move to lower bit widths when memory forces it, because quality loss grows and variant behaviour varies more.

Then verify with your own evaluation set rather than generic impressions. Keep a fixed prompt list, score task success and format validity, and pin the exact model file you validated so deployments are reproducible. If local memory remains the constraint, a hosted OpenAI-compatible API serves 30+ models without any quantization decisions on your side. See pricing for plans and start free with plugsky-micro and plugsky-lite.

Honest comparison

ConcernQ4Q8Check before deciding
MemoryRoughly half of Q8Roughly double Q4Model size versus available memory
QualitySmall task-dependent dropNear-losslessYour evaluation set
SpeedUsually faster where bandwidth-boundSlower per tokenHardware profile
Use whenMemory or speed is tightThe model fits with headroomBudget
Typical presetQ4_K_MQ8_0Runtime support

Frequently asked questions

Is Q4 much worse than Q8?

On most everyday tasks the difference is small and often invisible, but it depends on the model and workload. Tasks needing precise formatting or exact recall are more sensitive, so test before deciding.

How much memory does each level need?

As a rule, 4-bit weights are roughly half a gigabyte per billion parameters and 8-bit roughly one gigabyte, before KV cache and runtime overhead.

Why is Q4 sometimes faster?

Generation is often limited by memory bandwidth, so smaller weights mean less data moved per token. On compute-bound hardware the gap narrows.

Should I use an even lower bit width?

Below four bits, quality loss becomes more likely and variant behaviour varies more. Use it only when memory forces it and validate carefully.

Does quantization affect the KV cache?

No, cache precision is a separate setting. You can quantize weights and leave the cache at higher precision, or the reverse, depending on what is tight.

Do hosted models use quantization?

Provider-side serving may use optimized formats internally, but you do not manage them. A hosted endpoint sidesteps local memory limits entirely.

What is the safest default?

Start at Q4_K_M. If quality fails on your evaluation set and the model fits at Q8, move up; if it fits at Q8 from the start, use it.