Key facts
| Tool type | Free three-scenario deployment cost estimator |
| Scenarios | Managed cloud API, dedicated VPC, on-prem hardware |
| Inputs | GPU class, utilization hours, storage, egress, operations overhead |
| Outputs | Estimated cost per scenario and a break-even utilization point |
| Deployment options | Plugsky cloud, your VPC, on-prem and air-gapped |
| Companion tools | GPU capacity calculator and self-hosting break-even calculator |
| Models | 30+ models behind one OpenAI-compatible endpoint |
| Product status | Live |
TL;DR
- Utilization decides the winner: idle GPUs are the most expensive line item.
- Include operations, not just hardware: staffing and upgrades dominate small deployments.
- VPC deployments sit between convenience and control for regulated workloads.
- Edit every assumption; vendor estimates are only as good as your utilization curve.
- Compare scenarios per workload — different workloads belong in different homes.
How it works, step by step
- List workloads and their sensitivity: data residency, latency and isolation requirements.
- Estimate sustained and peak utilization in GPU-hours per day for each workload.
- Open the private LLM cost estimator and pick the GPU class and storage profile.
- Add operations overhead: engineering time, monitoring, upgrades and spares.
- Review the cloud, VPC and on-prem estimates side by side.
- Find the break-even utilization and compare it with your real curve.
- Assign each workload to a deployment model and re-run when volume changes.
Try it yourself
Open the private LLM cost estimator →
The three scenarios are different products
A managed API is a service: someone else owns capacity, uptime and upgrades, and you pay per usage or per plan. A dedicated VPC is a middle path: your network boundary, your keys and your logs, with managed operations inside a private environment. On-prem gives maximum control and maximum responsibility — hardware purchase, power, cooling, staff and refresh cycles. The estimator prices all three because the honest answer is usually workload-specific: regulated inference in a VPC, high-volume batch on owned hardware, and everything spiky on the managed API.
Utilization is the whole argument
Hardware economics are a utilization story. At low, spiky usage, GPUs sit idle and the effective cost per token climbs above any managed plan. At sustained utilization, the fixed cost amortises and on-prem can win — but only if you actually sustain it. Enter your real curve, including nights and weekends, rather than planning around peak capacity. The estimator's break-even point is the number to watch: if your utilization sits below it, managed deployment is cheaper; if it sits above it consistently, private infrastructure deserves a serious look.
Costs the hardware quote omits
The purchase order is the visible part. Add engineering time for deployment, upgrades and incident response; monitoring and logging infrastructure; redundancy for failover; and the cost of over-provisioning to absorb peaks. For regulated environments, add audit evidence and the compliance work that surrounds the deployment. These lines are why small on-prem pilots often cost more than expected. A three-scenario estimator forces them into the open, so the comparison is between total costs rather than between a hardware quote and a service invoice.
Honest comparison
| Factor | Managed cloud API | Dedicated VPC | On-prem hardware |
|---|---|---|---|
| Upfront cost | None | Low to medium | High, plus facilities |
| Scaling | Immediate | Fast, within provisioned capacity | Slow, tied to procurement |
| Operations burden | Vendor-owned | Shared | Fully yours |
| Data control | Region choice and contractual controls | Your network boundary, your keys | Maximum isolation |
| Best fit | Spiky or low utilization | Regulated steady workloads | High sustained utilization |
Frequently asked questions
Does the estimator include GPU purchase costs?
It lets you choose a GPU class and utilization profile so you can model hardware, rental or managed alternatives with the same assumptions.
What utilization makes on-prem cheaper?
That depends on your inputs — GPU class, power, staff and refresh cycle. The estimator computes a break-even point; compare it with your real utilization curve.
Why is a VPC different from the public API?
A VPC deployment puts inference inside your network boundary with your keys and logs, while managed cloud API is a multi-tenant service with region and contractual controls.
Are operations costs included?
You enter them. Staffing, monitoring, upgrades and redundancy are usually the lines that make small on-prem deployments uneconomic.
Can Plugsky deploy in a VPC or on-prem?
Yes. Enterprise deployments support your VPC, on-prem and air-gapped environments, with region selection for data residency.
Should every workload use the same deployment?
No. A common pattern is managed API for spiky workloads, VPC for regulated steady traffic, and on-prem for high-volume or air-gapped needs.
How accurate is the estimate?
It reflects the assumptions you enter and is meant for scenario comparison, not procurement. Replace defaults with your own numbers before deciding.
Is there a way to start without infrastructure?
Yes. The free plan includes 2 free AI models with no card, and 30+ models are available from the managed API.