Agents

What is the best way to self-host an AI agent?

Self-hosting an agent means running the model server, orchestration layer and tools inside your own infrastructure. Options range from single-machine runtimes for prototypes to vLLM-class servers on GPUs for production. You gain privacy, residency and cost control at scale, and take on GPU capacity planning, upgrades, observability and around-the-clock operations.

Key facts

Prototype runtimesOllama, llama.cpp and LM Studio classes of local runtimes
Production servingvLLM, TensorRT-LLM and TGI classes of GPU servers
Hardware driverModel size, quantization and context length determine VRAM needs
Hybrid optionKeep data local and send only necessary context to a managed model
Managed alternativeOn-prem and air-gapped Plugsky deployment with the same API
PrivacyFull data control when nothing leaves your network
MaintenanceYou own upgrades, capacity, monitoring, failover and security patching
StatusPlugsky agents and function calling are live in cloud and private deployments

TL;DR

  • Self-host when residency, air-gap or scale economics require it — not by default.
  • Match the runtime to the stage: single-machine for prototypes, GPU servers for production.
  • VRAM needs follow model size, quantization and context length together.
  • Budget for operations, not just hardware: upgrades, monitoring and failover are ongoing.
  • Hybrid keeps sensitive data local while a managed model handles reasoning.

How it works, step by step

  1. State the constraint that forces self-hosting: residency, air-gap, cost at scale or margin.
  2. Prototype locally with a single-machine runtime and a small quantized model.
  3. Measure the quality your task actually needs before sizing hardware.
  4. Choose a production server for concurrency, batching and observability.
  5. Size GPUs from model weights plus KV cache for your context length and concurrency.
  6. Build the surrounding platform: queue, sandboxed tools, tracing, evaluation and failover.
  7. Compare total cost with a managed or hybrid option before committing to a fleet.
1State theconstraint thatforces2Prototype locallywith asingle-machine3Measure the qualityyour task actuallyneeds before sizing4Choose a productionserver forconcurrency,5Size GPUs frommodel weights plusKV cache for your6Build thesurroundingplatform: queue,

Try it yourself

Open the self-hosting requirements tool →

What self-hosting actually involves

The model server is the visible part; the platform around it is the real work. A self-hosted agent needs orchestration, a tool sandbox, secrets management, tracing, evaluation, model versioning, capacity planning and failover. Running one agent on one machine is easy; running a reliable service on your own hardware is an operations commitment.

Privacy and residency are the usual drivers, and they are legitimate. Where self-hosting also wins is at steady, high utilization: if GPUs stay busy, the economics can beat usage-based pricing. Where it loses is elasticity — bursty demand means idle expensive hardware.

Choosing a stack by scale

  • Prototype: a single-machine runtime with a small quantized model, used to validate prompts and tools.
  • Small production: one or two GPU servers running a modern serving stack with batching.
  • Team production: multiple replicas behind a router, with model versioning and fallback.
  • Sovereign or air-gapped: the same serving layer inside an isolated network, with offline model distribution and no external dependencies.

Sizing starts with weights plus KV cache. Quantization reduces weight memory but can affect quality; larger context and more concurrent requests multiply cache memory. Measure with your own prompts rather than trusting rules of thumb.

The hybrid middle path

Many teams do not need to choose between fully local and fully managed. Sensitive data and tools stay in your network, while model calls that need frontier reasoning go to a managed endpoint; routine classification and rewriting run locally on a small model. This keeps the highest-risk data inside your control and the hardest reasoning on the best available models.

Plugsky supports this shape directly: the same OpenAI-compatible API runs in the shared cloud, in your VPC, on-prem and air-gapped, with region choice for residency, plus scoped keys, RBAC and audit logging. 30+ models behind one key make routing a configuration choice, so you can keep the local tier for private data and burst to managed capacity when needed. Chat, function calling, embeddings, RAG and agents are live; fine-tuning and batch endpoints are coming soon. Plans are on the live pricing page.

Honest comparison

OptionControlEffortBest for
Local single-machine runtimeHighLowPrototypes and demos
Self-hosted GPU serversHighestHighSteady high utilization, air-gap
Hybrid local plus managedHigh for sensitive dataMediumRegulated teams with frontier needs
Managed cloud APIProvider-definedLowestMost products and bursty demand
Private managed deploymentHigh, contract-backedLowEnterprises wanting residency without ops

Frequently asked questions

How much VRAM do I need?

It depends on model size, quantization, context length and concurrency. Weights plus KV cache define the floor; measure with your own prompts and add headroom for peaks.

Is self-hosting cheaper?

Only at high, steady utilization. Bursty workloads often cost less on managed or flat-rate plans because you do not pay for idle GPUs.

Can I self-host and use Plugsky together?

Yes. Keep sensitive data and tools local, and route model calls to Plugsky in the cloud, your VPC, on-prem or air-gapped depending on the workflow.

What about air-gapped environments?

They are possible but demanding: you need offline model distribution, local registries and a full observability stack. Plugsky's air-gapped deployment uses the same API, which reduces the application-side work.

Do local models support function calling?

Many modern open-weight models do, with varying reliability. Validate tool-call accuracy on your task set before making it the production path.

How do I keep a self-hosted agent secure?

Sandbox tools, scope credentials, restrict egress, log every action and patch the serving stack promptly. Self-hosting does not remove the need for agent-level controls.

What should I measure before committing?

Quality on your tasks, cost per completed task, GPU utilization and operational hours per month. Compare those against a managed or hybrid baseline.