Local AI

How do you run local AI on Windows?

On Windows, run local models with a native CUDA app such as LM Studio or Ollama, or inside WSL2 using the Windows GPU driver. Native apps are simpler and see the GPU directly; WSL2 gives a Linux toolchain with GPU passthrough. Either way, plan VRAM first, because Windows shares GPU memory with the desktop.

Key facts

Native runtimesLM Studio, Ollama and llama.cpp ship Windows builds with CUDA
WSL2 optionA Linux toolchain with GPU passthrough using the Windows driver
Driver ruleInstall the Windows GPU driver; do not add a second driver inside WSL
VRAM sharingDesktop composition, browsers and apps consume GPU memory
Model formatsGGUF for llama.cpp-family runtimes; Safetensors for GPU servers
Fallback pathPlugsky OpenAI-compatible API for tasks beyond local capacity
Endpoint statusChat, streaming, JSON mode, function calling and embeddings live

TL;DR

  • Native apps are the shortest path; WSL2 is for Linux-only tooling.
  • Install the GPU driver on Windows only, since WSL2 uses it through passthrough.
  • Leave VRAM headroom for the desktop, browsers and Windows itself.
  • GGUF covers most Windows runtimes, while vLLM-style servers prefer Linux.
  • Hybrid routing keeps local privacy with hosted capability on demand.

How it works, step by step

  1. Confirm your GPU model, VRAM and current NVIDIA or AMD driver version.
  2. Install a native runtime (LM Studio, Ollama or llama.cpp) and pull a 4-bit model.
  3. Verify GPU offload is enabled and check dedicated GPU memory usage.
  4. Optionally set up WSL2 using the driver already installed on Windows.
  5. Keep context modest, then raise it while watching VRAM.
  6. Test throughput with your real prompts and workload.
  7. Add an OpenAI-compatible hosted endpoint as fallback for larger models.
1Confirm your GPUmodel, VRAM andcurrent NVIDIA or2Install a nativeruntime (LM Studio,Ollama or3Verify GPU offloadis enabled andcheck dedicated GPU4Optionally set upWSL2 using thedriver already5Keep contextmodest, then raiseit while watching6Test throughputwith your realprompts and

Try it yourself

Open the Docker Compose generator for local AI →

Native Windows versus WSL2

Native applications are the fastest way to start. LM Studio gives you a desktop interface with a built-in local server; Ollama runs as a background service with a command line and REST API; llama.cpp offers the most control over quantization and offload. All three use CUDA on NVIDIA hardware and ROCm on supported AMD cards.

WSL2 matters when a tool only targets Linux, such as some GPU serving stacks and Docker-based pipelines. It runs a real Linux kernel with GPU passthrough, so the same commands from Linux guides work. The cost is extra disk, extra memory overhead and one more layer to debug.

Driver, VRAM and model gotchas

The most common Windows problem is driver confusion: install the GPU driver on Windows only. WSL2 uses that driver through passthrough, and adding a Linux driver inside the distribution can break acceleration. Keep the driver current, because runtime builds expect recent CUDA versions.

  • VRAM is shared. Windows, browsers and overlays take a slice before your model loads.
  • Task Manager is your gauge. Watch dedicated GPU memory, not shared memory.
  • GGUF is portable. Most Windows runtimes consume single-file GGUF models.
  • Servers prefer Linux. vLLM and similar stacks target Linux, so use WSL2 or a Linux host.

If a model fits with offload only, expect slow generation. A smaller fully GPU-resident model is usually the better daily choice.

Building a hybrid Windows setup

Windows machines are well suited to a hybrid pattern: keep chat, document work and offline tasks on the local runtime, and send heavy reasoning or long-context jobs to an OpenAI-compatible cloud endpoint. Because both sides speak the same request shape, the switch is configuration rather than a rewrite.

Plugsky serves 30+ models behind one OpenAI-compatible API with flat monthly self-serve plans and private deployment options. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; audio, image, moderation, files, batch and fine-tuning endpoints are coming soon. Check pricing and start free with plugsky-micro and plugsky-lite.

Honest comparison

ConcernNative Windows appWSL2 with GPU passthroughHosted Plugsky API
Setup effortInstaller plus a model downloadLinux environment and CUDA toolchainAPI key and a base URL
GPU accessDirect, through the Windows driverPassthrough using the Windows driverProvider-managed GPUs
ToolingWindows builds of popular runtimesLinux-only serving stacks availableFully managed
VRAM overheadDesktop and apps share GPU memorySimilar, plus WSL overheadNone locally
Best forDesktop chat and single-user workLinux-first development and servingScale, larger models and uptime

Frequently asked questions

Do I need WSL2 to run local LLMs on Windows?

No. LM Studio, Ollama and llama.cpp run natively with CUDA or ROCm. WSL2 helps when a tool is Linux-only, such as some server stacks.

Which GPU driver should I install for WSL2?

Install the standard Windows GPU driver only. WSL2 uses it through GPU passthrough, and installing a Linux driver inside can break acceleration.

Why is my GPU memory almost full before loading a model?

Windows, browsers and desktop composition use GPU memory. Close GPU-heavy apps and leave headroom, or reduce model size and context.

Can I run vLLM on Windows?

vLLM targets Linux. The practical route is WSL2 or a Linux host, or a managed OpenAI-compatible service if you need batching without Linux operations.

Is local AI on Windows private?

Inference stays on the machine, but the runtime and any extensions may call the network. Check telemetry settings and firewall rules.

How do I choose between native and WSL2?

Choose native for a desktop assistant or single-user experiments, and WSL2 when you need Linux tooling, Docker-based stacks or server-style serving.

How do I move a Windows pilot to production?

Containerise the service behind an OpenAI-compatible API and decide whether production runs on a Linux GPU server or a private hosted endpoint.