Key facts
| Billing unit | Requests, tokens, seats, committed spend or flat plan — clarify which |
| Commitments | Minimum annual spend, true-up rules and price-protection windows |
| Overage | Self-serve Plugsky plans have no per-token charges or overage fees |
| SLA | Published uptime target and service credits; see the SLA page |
| Residency | Region choice plus VPC, on-prem and air-gapped deployment options |
| Security | SSO, RBAC, audit logs, DPA and BYOK options for regulated buyers |
| Support | Named contacts, response targets and architecture reviews by tier |
| Product status | Enterprise deployment options are live; some specialist endpoints are coming soon |
TL;DR
- The headline rate is rarely the number that blows the budget — commitments and overage are.
- Self-serve Plugsky plans carry no per-token charges or overage fees.
- Residency, audit and exit terms are pricing terms in regulated industries.
- Get price protection and true-up rules in writing before the pilot ends.
- Compare five-year TCO including migration and operations, not year one only.
How it works, step by step
- Define the workload: models, monthly requests, peak concurrency and residency needs.
- Ask every vendor for a written pricing schedule covering the full contract term.
- Model three scenarios — base, growth and spike — including overage treatment.
- Price the non-token items: support tier, SSO, audit, DPA review and migration effort.
- Check SLA targets, service credits and the process for missed targets.
- Negotiate price protection, true-up windows and exit or data-portability terms.
- Run a paid pilot with production-like traffic before committing annual spend.
Try it yourself
Open the AI model pricing calculator →
Compare TCO, not the token rate
Two vendors can quote the same rate per million tokens and land in different budget lines. The difference lives in the surrounding terms: whether unused commitment rolls over, how overage is charged, what support costs, whether SSO and audit logs are included, and how much engineering time the integration consumes.
Build a spreadsheet with columns for platform fees, commitment, usage, overage, support, security review, migration and exit. A flat-rate plan collapses most of those columns into one predictable line, which is exactly why finance teams like it.
The procurement checklist
- Unit of billing — per token, per request, per seat, per plan or committed spend.
- Minimum commitment — annual floor, ramp schedule and what happens if you underuse.
- Overage — rate, alerts and whether there is a hard cap. Plugsky self-serve plans have neither per-token charges nor overage fees.
- SLA — uptime target, measurement window and service credits.
- Residency and deployment — region selection, VPC, on-prem and air-gapped options.
- Security and compliance — SSO, RBAC, audit logs, DPA and BYOK.
- Exit — data export, API compatibility and notice periods.
Why flat-rate pricing simplifies procurement
Variable billing creates three problems for large buyers: forecasts miss, budget owners hoard, and finance cannot approve a number that changes monthly. Flat monthly plans with unlimited fair-use usage remove per-token arithmetic and overage risk on self-serve, while enterprise agreements add committed capacity, negotiated limits and deployment options.
That does not make flat-rate automatically cheaper. It makes the number knowable, which shortens approval cycles and removes the end-of-quarter surprise that erodes trust in the AI programme.
Questions that expose hidden cost
- What exactly happens on the first day we exceed the committed volume?
- Are retries, failed calls and evaluation runs billed?
- Which support tier is required for production incidents, and what does it cost?
- Is SSO, audit logging and a signed DPA included or an upgrade?
- Can we deploy in our own VPC or on-prem, and at what price?
- What is the exit path if we migrate to another provider in year three?
Answers to those six questions usually move the TCO more than a rate negotiation.
Honest comparison
| Procurement factor | Plugsky | Typical enterprise token contract | Self-managed open source |
|---|---|---|---|
| Billing | Flat monthly plans; enterprise committed capacity | Committed spend with per-token overage | GPU capex plus platform ops |
| Overage risk | None on self-serve plans | True-ups and overage rates apply | Utilisation risk sits with you |
| Residency | Region choice, VPC, on-prem, air-gapped | Usually a limited region list | You control the hardware |
| SLA | Published SLA with service credits | Varies; often negotiated | You build your own |
| Security | SSO, RBAC, audit logs, DPA, BYOK options | Tiered or add-on | You build and audit |
| Exit | OpenAI-compatible API, data export | Provider-specific APIs | Portable but costly to run |
Frequently asked questions
What is the biggest hidden cost in enterprise AI contracts?
Overage on committed volume, followed by support tier upgrades and security add-ons. Ask what happens on the first day you exceed the commitment and whether retries and evaluation runs are billed.
How does Plugsky price enterprise deployments?
Enterprise agreements cover committed capacity, negotiated limits, SLA terms and deployment options including VPC, on-prem and air-gapped. See the live pricing page and talk to the team for a written schedule.
Are there overage fees on Plugsky self-serve plans?
No. Self-serve plans are flat monthly with unlimited fair-use usage, and there are no per-token charges or overage fees on those plans. Enterprise terms are negotiated separately.
Should procurement compare token rates across vendors?
Only as one input. Compare TCO across commitments, overage, support, residency, security review, migration and exit. A lower token rate with a rigid commitment can cost more overall.
What should an AI vendor risk assessment cover?
Data handling, residency, subprocessors, security controls, incident history, SLA performance, financial stability and exit or portability terms. Plugsky publishes SLA and legal pages and documents security controls.
How long should a pilot run before committing?
Long enough to include peak traffic and at least one full billing cycle. Run production-like volume so the committed tier matches reality rather than a demo workload.
Does residency change the price?
It can, because sovereign and air-gapped deployments use dedicated infrastructure. Ask for a written schedule that separates platform fees from deployment and support costs.
What proof of capacity should we request?
Ask for concurrency and throughput documentation, status history and a reference architecture for your region. Committed capacity should be specified in the contract, not assumed.