The Quantization Calculator compares model size across common quantization formats, including GGUF, GPTQ, AWQ and FP8, showing each format's size and its percentage of FP32. Enter a parameter count and the table updates. It is for engineers fitting models into limited GPU or CPU memory and for anyone deciding how much precision to trade for size. Results are estimates, since actual file sizes vary with the quantization recipe.
Quantization stores model weights at lower precision, such as 4-bit or 8-bit instead of 16-bit, reducing memory use and often increasing speed. The trade-off is a small quality loss that varies by task and recipe. It is the standard way to run larger models on consumer hardware.
GGUF is popular for CPU and mixed CPU/GPU inference with llama.cpp style runtimes. GPTQ and AWQ target GPU inference with different calibration approaches, and FP8 suits newer data-centre GPUs. Match the format to your runtime first, then compare quality.
Usually little at 8-bit, and modest loss at 4-bit for general chat and summarisation. Coding, maths and long-context reasoning degrade sooner. Always evaluate the quantized model on your own prompts instead of trusting a single benchmark number.
Canonical pricing and plans: plugsky.com/#sec-pricing · Terms · SLA · Docs