Key facts
| Entry GPU | 1× H100 80GB for ~18B models; 4× A100 40GB for 70B-class |
| System | 128GB RAM, 8TB NVMe, 25GbE internal fabric |
| Key custody | Hardware HSM for keys |
| Inference | vLLM, TensorRT-LLM or SGLang |
| API layer | OpenAI-compatible gateway (Plugsky or open source) |
| Observability | OpenTelemetry, Prometheus, Grafana |
| Vector store | pgvector, Qdrant or Pinecone |
| Product status | Live (Enterprise engagement) |
TL;DR
- Self-hosting is a stack, not just a model: serving, gateway, registry, observability and ops.
- One H100 80GB is a realistic starting point for ~18B models; 70B-class needs more GPUs.
- vLLM, TensorRT-LLM and SGLang are the common inference servers.
- An OpenAI-compatible gateway keeps application code portable.
- Plan for patching, model updates, monitoring and on-call from day one.
How it works, step by step
- Estimate concurrency, context length and model size for your workload.
- Size GPUs, RAM, NVMe and fabric against that estimate, not a template.
- Choose an inference server: vLLM, TensorRT-LLM or SGLang.
- Put an OpenAI-compatible gateway in front so clients stay portable.
- Add observability, a model registry and HSM-based key custody.
- Define the patching, update and on-call process before go-live.
Try it yourself
Open the GPU capacity calculator →
What self-hosting really involves
Running a model on your own hardware is the easy part. A production on-prem LLM also needs an inference server, an API gateway that speaks the OpenAI protocol, a model registry with signed artifacts, observability, security patching and a team on call when something breaks. Teams that budget only for GPUs usually discover this in month two.
The payoff is control: your data never leaves the network, you choose exactly which model version runs, and you are not exposed to a vendor's deprecation schedule. That control comes with operational ownership.
Reference hardware and serving stack
The reference build starts with a single H100 80GB for approximately 18B-parameter models, or four A100 40GB GPUs for 70B-class models, plus 128GB RAM, 8TB NVMe and a 25GbE internal fabric. A hardware HSM holds keys. Larger deployments move to multi-node clusters with H100, H200 or MI300X pools, and CPU-only inference is possible for small models where latency is not critical.
- Inference: vLLM, TensorRT-LLM or SGLang.
- Gateway: Plugsky or an open-source OpenAI-compatible proxy.
- Observability: OpenTelemetry, Prometheus and Grafana.
- Vector store: pgvector, Qdrant or Pinecone.
Sizing your cluster
Three numbers drive sizing: concurrent requests, context length and model size. A model that fits comfortably in GPU memory at short context can fail at long context once the KV cache grows, so size for your real context distribution rather than the maximum the model advertises. Throughput also depends on batching — a serving engine that batches well can double effective capacity on the same hardware.
Use a calculator to get a first estimate, then benchmark with your own prompts. Published benchmarks rarely match production traffic mixes.
Ops, updates and security
Plan the operational model before go-live: how often you patch the OS and serving stack, how model weights are signed and staged, who is on call, and how upgrades roll back. Keep the previous model version available so a regression can be reverted quickly. Audit logging, key custody and network segmentation should be designed with the same seriousness as in a cloud deployment — an on-prem stack is not automatically more secure just because it is private.
If operating this stack is not your competitive advantage, compare it honestly against a managed private deployment before committing the headcount.
Honest comparison
| Factor | On-prem self-hosted | Plugsky VPC deployment | Plugsky public cloud |
|---|---|---|---|
| Hardware | You buy and run GPUs | Runs in your cloud account | Managed by Plugsky |
| Data path | Entirely inside your network | Inside your VPC | Region-locked data plane |
| Model updates | You stage and apply | Signed bundles | Continuous |
| Ops burden | Highest | Moderate | Lowest |
| Latency control | Full control | Good | Region-dependent |
| Time to first inference | Weeks to months | Days to weeks | Minutes |
Frequently asked questions
What hardware do I need for an 18B model?
The reference build uses one H100 80GB with 128GB RAM, 8TB NVMe and a 25GbE fabric; add an HSM for key custody.
What about 70B-class models?
The reference starts at four A100 40GB GPUs, though exact sizing depends on context length, concurrency and quantisation.
Which inference server should I use?
vLLM, TensorRT-LLM and SGLang are the common choices. Pick based on your hardware, model format and batching needs, then benchmark.
Do I need an API gateway?
Yes, if you want portability. An OpenAI-compatible gateway lets the same client code run against on-prem, VPC or cloud deployments.
How do model updates work?
Weights are staged through a registry as signed artifacts and applied with rollback to the previous version; the update process should be rehearsed before go-live.
Is on-prem always more secure?
No. It removes vendor access, but you inherit patching, hardening and monitoring. Security depends on how well the stack is operated.
Can Plugsky manage an on-prem deployment?
Yes. Managed AI Ops and Enterprise support cover on-prem and air-gapped environments with defined service levels.
Plugsky (2026). “On-Prem LLM Deployment — Architecture and Sizing”. Plugsky. Available at: https://plugsky.com/solutions/on-prem-llm (last updated 2026-09-25).