Key facts
| Code model families | Qwen Coder, DeepSeek Coder, Codestral plus general Llama, Mistral and gpt-oss models |
| Common sizes | 1.5B to 34B parameters; 7B-14B at 4-bit is the practical sweet spot |
| Memory guide | A 7B model at 4-bit needs roughly 4-5 GB of weights before KV cache |
| Autocomplete | Fill-in-the-middle training improves inline completions in editors |
| Context length | Coder models commonly support 32k tokens or more; KV cache grows with context |
| Editor integrations | Continue, Cline and similar tools accept OpenAI-compatible endpoints |
| Cloud fallback | Plugsky exposes coding-capable models through an OpenAI-compatible API |
TL;DR
- Code-specialised models beat general models at the same size for completion and refactoring.
- Use fill-in-the-middle support if you want inline editor autocomplete, not just chat.
- 8 GB of VRAM runs a 7B coder model at 4-bit; 16-24 GB handles larger models or longer context.
- Keep the endpoint OpenAI-compatible so your editor plugin works with local and hosted models.
- Split the workflow: local models for completions and edits, hosted models for hard debugging.
How it works, step by step
- Decide the workflow you need: inline completion, chat in the editor, or repository-wide edits.
- Check your memory budget and pick a coder model that fits with room for context.
- Serve the model with Ollama, llama.cpp or LM Studio and confirm the OpenAI-compatible endpoint.
- Point your editor plugin at the local endpoint and set a completion timeout that fits your machine.
- Test on your own language and framework, including one long file and one refactor task.
- Tune context and stop sequences to keep completions short and latency low.
- Add a hosted fallback for tasks that the local model fails or cannot fit in context.
Original data
Try it yourself
Open the best model for coding selector →
What makes a coding model good locally
Three capabilities matter. Completion quality covers single-line and block suggestions, which is mostly a function of code-specific training and fill-in-the-middle support. Instruction following matters for chat-style requests such as explaining, refactoring or writing tests. Context length decides how much of a file or repository the model can see at once.
Size changes the trade-off. A 7B coder model at 4-bit is fast enough for inline completion and fits modest hardware. A 14B or 32B model produces better multi-file reasoning but needs more memory and time per token. Most developers get better results routing short completions to a small model and longer reasoning tasks to a larger one, local or hosted.
Matching models to hardware
Weights are the first budget line: parameters times bits per weight. At 4-bit, a 7B model needs roughly 4-5 GB, a 14B about 8-9 GB and a 32B around 17-19 GB, before KV cache and runtime overhead. Context then adds on top, which is why a 32B model may technically load but leave no room for the long files you wanted it to read.
- 8 GB VRAM: 7B coder models at 4-bit with modest context.
- 12-16 GB: 7B at higher precision or 14B at 4-bit.
- 24 GB: 32B at 4-bit or 14B at 8-bit.
- Apple Silicon: unified memory raises the ceiling, bandwidth sets the speed.
Quantize to Q4 or Q5 for most work. Going below 4-bit saves memory but degrades code correctness in ways that are hard to notice until tests fail.
Editor integration and hybrid routing
Editor plugins do not need a bespoke integration: they need an OpenAI-compatible endpoint and a model name. Run your local server, point the plugin at it, and set conservative timeouts so completions do not block typing. Keep an eye on context size, because editors often send large file prefixes that quietly consume your memory budget.
Local models stumble on sprawling refactors or unfamiliar APIs. A hybrid setup handles that: completions and routine edits stay local, while hard tasks go to a hosted model. Plugsky serves coding-capable models alongside chat, streaming, JSON mode and function calling through one OpenAI-compatible API covering 30+ models, so switching is a base URL change. The free plan includes plugsky-micro and plugsky-lite, and a 14-day full-access trial covers evaluation. See the live pricing page before choosing a tier.
Honest comparison
| Setup | Model size | Best for | Memory | Speed |
|---|---|---|---|---|
| CPU with llama.cpp | 1.5B-7B at 4-bit | Occasional completion | System RAM | Slow |
| 8 GB GPU | 7B-8B at 4-bit | Inline completion | Fits with modest context | Fast |
| 16 GB GPU | 14B at 4-bit | Chat plus edits | Comfortable | Fast |
| 24 GB GPU | 32B at 4-bit | Multi-file reasoning | Needs context budget | Moderate |
| Hybrid with Plugsky | Hosted models | Hard debugging and refactors | None local | Network-bound |
Frequently asked questions
Can a local model replace GitHub Copilot?
For inline completion on common languages, a good 7B-14B coder model gets close enough for many workflows. For complex multi-file reasoning, hosted frontier models still have an edge.
Which local model is best for coding?
Start with a code-specialised family such as Qwen Coder or DeepSeek Coder at 7B-14B, then test against general models like Llama, Mistral or gpt-oss on your own codebase.
How much VRAM do I need for a local coding model?
8 GB runs a 7B model at 4-bit with limited context. 16 GB is comfortable for 14B models, and 24 GB allows 32B models at low-bit quantization.
Do local coding models support fill-in-the-middle?
Many code-specialised models are trained for fill-in-the-middle, which improves inline completions. Verify your editor plugin sends the right prompt format.
Is my code private when I run locally?
Yes, prompts and completions stay on your machine. If you add a cloud fallback, only the prompts you route there leave your network.
Will a quantized model write worse code?
At 4-bit K-quant levels the difference is usually small. Below 4-bit, expect more syntax slips and logic errors, especially on longer generations.
How do I handle very large repositories?
Use retrieval to send only relevant files and symbols to the model, keep chunking consistent, and reserve long-context requests for tasks that truly need them.