Check if any LLM fits on your GPU. Enter model size, quantization, context length, and VRAM.
The calculator estimates total VRAM needed including model weights, KV cache, and overhead. If the total fits within your GPU's VRAM, the model can run. If not, try a smaller quantization, shorter context, or smaller model.
| GPU | VRAM | Max 7B Q4 | Max 70B Q4 |
|---|---|---|---|
| RTX 4060 | 8 GB | โ Yes | โ No |
| RTX 4090 | 24 GB | โ Yes | โ ๏ธ 4-bit |
| A100 | 80 GB | โ Yes | โ Yes |