Select your hardware and get a recommendation.
The Local Model Recommender suggests open models that match your GPU's memory. Choose a VRAM tier from 4 GB to 80 GB and your intended use case, and the tool returns model and quantization combinations that typically fit, from small 2 to 3 billion parameter models on 4 GB up to 70B models in full precision on large cards. It is for developers planning local inference before downloading weights.
It depends on parameter count and precision. The recommender's tiers show 7 to 8 billion parameter models at 4-bit on 8 GB, 14B at 4-bit on 12 GB, and 70B at 4-bit on 24 GB, with larger cards allowing higher precision and longer context.
Yes, but usually modestly. 8-bit is close to full precision for most tasks, while 4-bit is acceptable for chat and summarization and shows more degradation on complex reasoning or code. Test the quantized build on your own tasks.
Check license terms for commercial use, the context length the build supports, and whether your runtime offers a prebuilt binary for the quantization format. Also confirm the download fits your disk and leaves VRAM headroom for the KV cache.
Canonical pricing and plans: plugsky.com/#sec-pricing · Terms · SLA · Docs