Key facts
| Prototype runtimes | Ollama, llama.cpp and LM Studio classes of local runtimes |
| Production serving | vLLM, TensorRT-LLM and TGI classes of GPU servers |
| Hardware driver | Model size, quantization and context length determine VRAM needs |
| Hybrid option | Keep data local and send only necessary context to a managed model |
| Managed alternative | On-prem and air-gapped Plugsky deployment with the same API |
| Privacy | Full data control when nothing leaves your network |
| Maintenance | You own upgrades, capacity, monitoring, failover and security patching |
| Status | Plugsky agents and function calling are live in cloud and private deployments |
TL;DR
- Self-host when residency, air-gap or scale economics require it — not by default.
- Match the runtime to the stage: single-machine for prototypes, GPU servers for production.
- VRAM needs follow model size, quantization and context length together.
- Budget for operations, not just hardware: upgrades, monitoring and failover are ongoing.
- Hybrid keeps sensitive data local while a managed model handles reasoning.
How it works, step by step
- State the constraint that forces self-hosting: residency, air-gap, cost at scale or margin.
- Prototype locally with a single-machine runtime and a small quantized model.
- Measure the quality your task actually needs before sizing hardware.
- Choose a production server for concurrency, batching and observability.
- Size GPUs from model weights plus KV cache for your context length and concurrency.
- Build the surrounding platform: queue, sandboxed tools, tracing, evaluation and failover.
- Compare total cost with a managed or hybrid option before committing to a fleet.
Try it yourself
Open the self-hosting requirements tool →
What self-hosting actually involves
The model server is the visible part; the platform around it is the real work. A self-hosted agent needs orchestration, a tool sandbox, secrets management, tracing, evaluation, model versioning, capacity planning and failover. Running one agent on one machine is easy; running a reliable service on your own hardware is an operations commitment.
Privacy and residency are the usual drivers, and they are legitimate. Where self-hosting also wins is at steady, high utilization: if GPUs stay busy, the economics can beat usage-based pricing. Where it loses is elasticity — bursty demand means idle expensive hardware.
Choosing a stack by scale
- Prototype: a single-machine runtime with a small quantized model, used to validate prompts and tools.
- Small production: one or two GPU servers running a modern serving stack with batching.
- Team production: multiple replicas behind a router, with model versioning and fallback.
- Sovereign or air-gapped: the same serving layer inside an isolated network, with offline model distribution and no external dependencies.
Sizing starts with weights plus KV cache. Quantization reduces weight memory but can affect quality; larger context and more concurrent requests multiply cache memory. Measure with your own prompts rather than trusting rules of thumb.
The hybrid middle path
Many teams do not need to choose between fully local and fully managed. Sensitive data and tools stay in your network, while model calls that need frontier reasoning go to a managed endpoint; routine classification and rewriting run locally on a small model. This keeps the highest-risk data inside your control and the hardest reasoning on the best available models.
Plugsky supports this shape directly: the same OpenAI-compatible API runs in the shared cloud, in your VPC, on-prem and air-gapped, with region choice for residency, plus scoped keys, RBAC and audit logging. 30+ models behind one key make routing a configuration choice, so you can keep the local tier for private data and burst to managed capacity when needed. Chat, function calling, embeddings, RAG and agents are live; fine-tuning and batch endpoints are coming soon. Plans are on the live pricing page.
Honest comparison
| Option | Control | Effort | Best for |
|---|---|---|---|
| Local single-machine runtime | High | Low | Prototypes and demos |
| Self-hosted GPU servers | Highest | High | Steady high utilization, air-gap |
| Hybrid local plus managed | High for sensitive data | Medium | Regulated teams with frontier needs |
| Managed cloud API | Provider-defined | Lowest | Most products and bursty demand |
| Private managed deployment | High, contract-backed | Low | Enterprises wanting residency without ops |
Frequently asked questions
How much VRAM do I need?
It depends on model size, quantization, context length and concurrency. Weights plus KV cache define the floor; measure with your own prompts and add headroom for peaks.
Is self-hosting cheaper?
Only at high, steady utilization. Bursty workloads often cost less on managed or flat-rate plans because you do not pay for idle GPUs.
Can I self-host and use Plugsky together?
Yes. Keep sensitive data and tools local, and route model calls to Plugsky in the cloud, your VPC, on-prem or air-gapped depending on the workflow.
What about air-gapped environments?
They are possible but demanding: you need offline model distribution, local registries and a full observability stack. Plugsky's air-gapped deployment uses the same API, which reduces the application-side work.
Do local models support function calling?
Many modern open-weight models do, with varying reliability. Validate tool-call accuracy on your task set before making it the production path.
How do I keep a self-hosted agent secure?
Sandbox tools, scope credentials, restrict egress, log every action and patch the serving stack promptly. Self-hosting does not remove the need for agent-level controls.
What should I measure before committing?
Quality on your tasks, cost per completed task, GPU utilization and operational hours per month. Compare those against a managed or hybrid baseline.