Key facts
| Bit width | Both are 8-bit formats, about half the memory of FP16 |
| FP8 variants | E4M3 for weights and activations, E5M2 where wider range is needed |
| Hardware | FP8 acceleration requires newer data-centre and consumer GPUs; INT8 support is broader |
| Accuracy technique | INT8 often uses per-channel scaling or outlier-aware methods; FP8 uses scaling factors |
| KV cache | Both formats can quantize the KV cache to cut long-context memory |
| Runtime support | vLLM and modern inference stacks support FP8 and INT8 checkpoints |
| Cloud option | Plugsky is OpenAI-compatible with 30+ models, so quantization stays a serving concern |
TL;DR
- Choose FP8 when your GPU supports it and you want simpler scaling with wide dynamic range.
- Choose INT8 when hardware support is broader or tooling already targets integer kernels.
- Both roughly halve weight memory versus FP16 with modest quality loss in most cases.
- Quantizing the KV cache is often the bigger win at long context.
- Verify quality on your own evaluation set; format comparisons are workload-specific.
How it works, step by step
- Check which 8-bit formats your GPU and runtime actually support.
- Establish an FP16 or BF16 baseline on your evaluation set.
- Quantize to FP8 or INT8 using a supported pipeline and note the checkpoint format.
- Compare quality, latency and memory against the baseline at the same context length.
- Enable KV cache quantization separately and re-measure long-context behaviour.
- Pick the format that meets your quality bar at the capacity you need.
- Revisit the choice when you change GPU generation or serving stack.
Original data
Try it yourself
Open the quantization calculator →
How the two formats represent numbers
INT8 stores integers from -128 to 127 and maps model values into that range using a scale, sometimes a zero point. The mapping is efficient and predictable, but a few large outlier values can consume most of the range unless the quantizer handles them with per-channel scales or outlier-aware methods.
FP8 keeps a floating-point layout. E4M3 gives three mantissa bits and four exponent bits, which preserves dynamic range better than INT8 at the cost of precision within each value. That range advantage is why FP8 often needs less calibration: values that fall outside a fixed integer range are representable without special handling.
Hardware, runtimes and KV cache
Hardware decides first. FP8 tensor-core support arrived with recent GPU generations, so older cards must fall back to emulation or accumulate in higher precision, which erodes the speed benefit. INT8 kernels have existed longer and are available on a wider range of accelerators, including many consumer GPUs and some edge devices.
Runtime support is the second filter. Modern GPU serving stacks support FP8 checkpoints and FP8 KV cache in addition to INT8 paths.
- Confirm your serving stack lists the exact format and model architecture you plan to run.
- Check whether quantization is applied at load time or requires a pre-quantized checkpoint.
- Test KV cache quantization independently, since it affects context capacity, not weights.
Choosing between them
Quality differences between well-implemented FP8 and INT8 are usually small, and neither wins universally. FP8 tends to be simpler to apply on new hardware and is often the default in modern serving recipes. INT8 remains attractive when your hardware lacks FP8 support, when existing tooling already produces int8 checkpoints, or when you need the widest possible deployment compatibility.
Treat the choice as a serving decision, not a product decision. Keep evaluation prompts and metrics fixed, quantify the memory and throughput gain, and confirm the quality bar still holds. If you serve through a hosted API, the format question disappears: Plugsky exposes an OpenAI-compatible endpoint with 30+ models and live chat, streaming, JSON mode, function calling, embeddings, RAG and agents, while audio, image, moderation, batch and fine-tuning endpoints are coming soon. See the live pricing page for plan options.
Honest comparison
| Factor | FP8 | INT8 | FP16 baseline |
|---|---|---|---|
| Representation | Floating point, limited mantissa | Integer with scale | Floating point |
| Dynamic range | Wide | Narrow without outlier handling | Very wide |
| Hardware support | Newer GPUs | Broad | Universal |
| Memory vs FP16 | About half | About half | Reference |
| Typical use | Modern GPU serving and KV cache | Compatibility-first deployments | Quality reference |
Frequently asked questions
Does FP8 lose more quality than INT8?
Not inherently. FP8's wider range often simplifies scaling, while INT8 quality depends heavily on how outliers are handled. Compare both on your own evaluation set.
Which GPUs support FP8?
FP8 tensor-core acceleration requires recent GPU generations. Older cards can emulate FP8, but the throughput advantage is limited without native support.
Can I quantize the KV cache separately?
Yes, and it is often worthwhile for long context. KV cache quantization reduces memory at longer sequence lengths and can be combined with 16-bit weights.
Is INT8 still relevant in 2026?
Yes. INT8 kernels run on a wider range of hardware and remain common in compatibility-focused deployments.
Do I need calibration data?
INT8 pipelines commonly use calibration or scaling analysis. FP8 typically needs scaling factors too, though its dynamic range reduces sensitivity to outliers.
How do I measure the impact?
Fix an evaluation set and a context length, run the FP16 baseline, then compare quality, tokens per second and memory for each 8-bit format.
Should I use 8-bit or 4-bit quantization?
8-bit preserves more quality for a given model; 4-bit saves more memory and lets a larger model fit. At tight memory budgets, a larger model at 4-bit often beats a smaller one at 8-bit.