Key facts
| Self-serve pricing | Flat monthly with unlimited fair-use usage |
| Per-token charges | None on self-serve plans |
| Overage fees | None on self-serve plans |
| Enterprise pricing | Quoted per agreement for dedicated deployments |
| Free tier | 2 free AI models (plugsky-micro, plugsky-lite), no card |
| Trial | 14-day full-access trial available |
| Fair use | Rate and abuse limits apply to protect shared capacity |
| Product status | Live |
TL;DR
- Self-serve plans are flat monthly, not metered per token.
- No overage fees or end-of-month token surprises on self-serve.
- Fair-use limits still apply — flat does not mean unlimited abuse.
- Dedicated and enterprise deployments are quoted in an agreement.
- Usage analytics stay available so you can forecast capacity honestly.
How it works, step by step
- Estimate your request volume and average prompt size to understand real usage.
- Pick a self-serve plan sized for that usage rather than per-token forecasts.
- Track dashboard usage weekly to catch unusual growth early.
- Set internal budgets and quotas per team or product surface.
- Talk to the team before sustained peaks that could hit fair-use limits.
- Move to a dedicated or enterprise agreement when scale and guarantees demand it.
Try it yourself
Open the LLM cost calculator →
Why flat pricing changes planning
Per-token billing makes cost a function of usage, which makes forecasting a function of product behaviour — hard for teams shipping features. Flat monthly pricing moves cost into a fixed line item you can budget, and it removes the failure mode where a runaway loop or a viral feature produces a bill instead of an alert. The trade-off is a usage policy: flat plans rely on fair use, and providers enforce it so one customer cannot consume shared capacity without limit. For finance teams the benefit is simpler still: one recurring line item per plan instead of a forecast tied to product usage, which makes budget reviews and vendor comparisons straightforward.
What fair use means in practice
Fair use is about protecting shared capacity, not about hidden metering. Expect a few common-sense limits: requests are rate-limited to keep latency stable, abusive or automated patterns can be throttled, and sustained load far beyond a plan's design may need a conversation. Importantly, exceeding those limits leads to throttling or an upgrade discussion — not a variable bill. For applications with hard throughput requirements, an enterprise agreement with defined capacity is the honest answer rather than relying on soft limits.
What we do and what we do not do
What we do: publish flat self-serve plans, keep dashboards so usage is visible, and quote dedicated capacity transparently in an agreement. What we do not do: meter self-serve usage per token, invent overage charges, or market an unlimited plan that silently is not one — if a workload sits outside fair use, we say so. Read the terms for the usage policy and the live pricing page for current plans and contact paths.
Honest comparison
| Aspect | Plugsky self-serve | Typical per-token API | Plugsky enterprise |
|---|---|---|---|
| Billing model | Flat monthly | Per million tokens | Quoted agreement |
| Cost predictability | Fixed line item | Varies with usage | Contracted |
| Overage risk | None | Can be significant | Defined in contract |
| Abuse protection | Fair-use limits | Rate limits and spend caps | Dedicated capacity |
| Capacity guarantees | Shared | Varies by tier | Committed |
| Best fit | Product teams and startups | Sporadic or experimental use | High and regulated scale |
Frequently asked questions
Is there really no per-token charge?
Correct — self-serve plans are flat monthly with unlimited fair-use usage. There are no per-token charges and no overage fees on self-serve plans.
What happens if I exceed fair use?
You may be rate-limited or asked to move to a higher plan or dedicated capacity. You will not receive a surprise per-token bill.
Are all models included in the flat price?
The self-serve plans expose the model catalogue under fair use; check the live pricing page for plan details and any model tiers.
Why do enterprise deployments get quoted?
Dedicated, VPC, on-prem and air-gapped deployments carry infrastructure and support commitments, so they are priced in an agreement rather than a flat self-serve plan.
Can I get a spend cap or budget alert?
Usage is visible in the dashboard and plans are fixed-price, which removes the need for token spend caps. For internal controls, set quotas per team or product surface.
Does the free plan include the 14-day trial?
The free plan includes two models with no card, and a separate 14-day full-access trial lets you evaluate stronger models before choosing a plan.
Where do I see current prices?
On the live pricing page, linked from the main navigation — prices are kept there rather than hard-coded in articles.