Key facts
| Native runtimes | LM Studio, Ollama and llama.cpp ship Windows builds with CUDA |
| WSL2 option | A Linux toolchain with GPU passthrough using the Windows driver |
| Driver rule | Install the Windows GPU driver; do not add a second driver inside WSL |
| VRAM sharing | Desktop composition, browsers and apps consume GPU memory |
| Model formats | GGUF for llama.cpp-family runtimes; Safetensors for GPU servers |
| Fallback path | Plugsky OpenAI-compatible API for tasks beyond local capacity |
| Endpoint status | Chat, streaming, JSON mode, function calling and embeddings live |
TL;DR
- Native apps are the shortest path; WSL2 is for Linux-only tooling.
- Install the GPU driver on Windows only, since WSL2 uses it through passthrough.
- Leave VRAM headroom for the desktop, browsers and Windows itself.
- GGUF covers most Windows runtimes, while vLLM-style servers prefer Linux.
- Hybrid routing keeps local privacy with hosted capability on demand.
How it works, step by step
- Confirm your GPU model, VRAM and current NVIDIA or AMD driver version.
- Install a native runtime (LM Studio, Ollama or llama.cpp) and pull a 4-bit model.
- Verify GPU offload is enabled and check dedicated GPU memory usage.
- Optionally set up WSL2 using the driver already installed on Windows.
- Keep context modest, then raise it while watching VRAM.
- Test throughput with your real prompts and workload.
- Add an OpenAI-compatible hosted endpoint as fallback for larger models.
Try it yourself
Open the Docker Compose generator for local AI →
Native Windows versus WSL2
Native applications are the fastest way to start. LM Studio gives you a desktop interface with a built-in local server; Ollama runs as a background service with a command line and REST API; llama.cpp offers the most control over quantization and offload. All three use CUDA on NVIDIA hardware and ROCm on supported AMD cards.
WSL2 matters when a tool only targets Linux, such as some GPU serving stacks and Docker-based pipelines. It runs a real Linux kernel with GPU passthrough, so the same commands from Linux guides work. The cost is extra disk, extra memory overhead and one more layer to debug.
Driver, VRAM and model gotchas
The most common Windows problem is driver confusion: install the GPU driver on Windows only. WSL2 uses that driver through passthrough, and adding a Linux driver inside the distribution can break acceleration. Keep the driver current, because runtime builds expect recent CUDA versions.
- VRAM is shared. Windows, browsers and overlays take a slice before your model loads.
- Task Manager is your gauge. Watch dedicated GPU memory, not shared memory.
- GGUF is portable. Most Windows runtimes consume single-file GGUF models.
- Servers prefer Linux. vLLM and similar stacks target Linux, so use WSL2 or a Linux host.
If a model fits with offload only, expect slow generation. A smaller fully GPU-resident model is usually the better daily choice.
Building a hybrid Windows setup
Windows machines are well suited to a hybrid pattern: keep chat, document work and offline tasks on the local runtime, and send heavy reasoning or long-context jobs to an OpenAI-compatible cloud endpoint. Because both sides speak the same request shape, the switch is configuration rather than a rewrite.
Plugsky serves 30+ models behind one OpenAI-compatible API with flat monthly self-serve plans and private deployment options. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; audio, image, moderation, files, batch and fine-tuning endpoints are coming soon. Check pricing and start free with plugsky-micro and plugsky-lite.
Honest comparison
| Concern | Native Windows app | WSL2 with GPU passthrough | Hosted Plugsky API |
|---|---|---|---|
| Setup effort | Installer plus a model download | Linux environment and CUDA toolchain | API key and a base URL |
| GPU access | Direct, through the Windows driver | Passthrough using the Windows driver | Provider-managed GPUs |
| Tooling | Windows builds of popular runtimes | Linux-only serving stacks available | Fully managed |
| VRAM overhead | Desktop and apps share GPU memory | Similar, plus WSL overhead | None locally |
| Best for | Desktop chat and single-user work | Linux-first development and serving | Scale, larger models and uptime |
Frequently asked questions
Do I need WSL2 to run local LLMs on Windows?
No. LM Studio, Ollama and llama.cpp run natively with CUDA or ROCm. WSL2 helps when a tool is Linux-only, such as some server stacks.
Which GPU driver should I install for WSL2?
Install the standard Windows GPU driver only. WSL2 uses it through GPU passthrough, and installing a Linux driver inside can break acceleration.
Why is my GPU memory almost full before loading a model?
Windows, browsers and desktop composition use GPU memory. Close GPU-heavy apps and leave headroom, or reduce model size and context.
Can I run vLLM on Windows?
vLLM targets Linux. The practical route is WSL2 or a Linux host, or a managed OpenAI-compatible service if you need batching without Linux operations.
Is local AI on Windows private?
Inference stays on the machine, but the runtime and any extensions may call the network. Check telemetry settings and firewall rules.
How do I choose between native and WSL2?
Choose native for a desktop assistant or single-user experiments, and WSL2 when you need Linux tooling, Docker-based stacks or server-style serving.
How do I move a Windows pilot to production?
Containerise the service behind an OpenAI-compatible API and decide whether production runs on a Linux GPU server or a private hosted endpoint.