Local Model Recommender

Select your hardware and get a recommendation.

Your hardware

Best model for your setup

What the Local Model Recommender does

The Local Model Recommender suggests open models that match your GPU's memory. Choose a VRAM tier from 4 GB to 80 GB and your intended use case, and the tool returns model and quantization combinations that typically fit, from small 2 to 3 billion parameter models on 4 GB up to 70B models in full precision on large cards. It is for developers planning local inference before downloading weights.

How to use it

  1. Select your GPU's VRAM tier.
  2. Select your intended use case, such as chat, coding, RAG, or vision.
  3. Read the recommended models and their quantization levels.
  4. Confirm a specific model with the GPU Fit Checker before downloading.
  5. Try the model in your local runtime and compare output against your needs.

FAQ

How much VRAM do I need for a local LLM?

It depends on parameter count and precision. The recommender's tiers show 7 to 8 billion parameter models at 4-bit on 8 GB, 14B at 4-bit on 12 GB, and 70B at 4-bit on 24 GB, with larger cards allowing higher precision and longer context.

Does quantization hurt quality?

Yes, but usually modestly. 8-bit is close to full precision for most tasks, while 4-bit is acceptable for chat and summarization and shows more degradation on complex reasoning or code. Test the quantized build on your own tasks.

What should I check before downloading a model?

Check license terms for commercial use, the context length the build supports, and whether your runtime offers a prebuilt binary for the quantization format. Also confirm the download fits your disk and leaves VRAM headroom for the KV cache.

Start Free →

Canonical pricing and plans: plugsky.com/#sec-pricing · Terms · SLA · Docs

Related

Best local LLMs for 8 GB, 16 GB and 24 GB

What is a local LLM?

Best local AI apps