Key facts
| Cost unit | Cost per completed task, not cost per token |
| Main drivers | Turns multiplied by tokens, plus tools, retrieval and retries |
| Reduction levers | Model routing, caching, context trimming, turn caps |
| Self-serve billing | Flat monthly plans with unlimited fair-use usage |
| Free tier | plugsky-micro and plugsky-lite free, no card required |
| Trial | 14-day full-access trial for paid models |
| Visibility | Usage dashboard shows tokens, spend and model mix |
| Current amounts | Live prices are published on the Plugsky pricing page |
TL;DR
- Budget per completed task, because a cheap model that loops can cost more than a strong one that finishes.
- Turn count is the biggest multiplier: cap it and count tools per run.
- Route routine steps to small models and reserve frontier models for planning.
- Cache stable context and trim tool results to cut tokens on every turn.
- Flat monthly plans remove per-token forecasting; check the live pricing page.
How it works, step by step
- Instrument the agent to record tokens, turns, tool calls and retries per completed task.
- Calculate a baseline cost per task for each workflow before optimising anything.
- Route classification and rewriting to the cheapest model that passes your evaluations.
- Cache stable context such as policies and schemas, and trim verbose tool results.
- Set turn and tool-call caps so a confused run stops early instead of billing on.
- Re-run the task set after each change to confirm quality held while cost fell.
- Compare usage against flat plan tiers on the live pricing page and move when the math favours it.
Try it yourself
Open the agent run cost calculator →
Why agent pricing is different
A single chat request costs one prompt and one completion. An agent might make ten model calls, run eight tools, retrieve documents twice, fail once and retry. Every one of those steps consumes something: tokens, API calls, compute, database reads or third-party fees. Pricing an agent per token hides that structure; pricing it per completed task exposes it.
That reframing changes optimisation too. A slightly more expensive model that finishes in three turns can beat a cheaper model that wanders for twelve. Measure outcomes, not unit prices.
The cost stack
- Model tokens: turns multiplied by input and output tokens, including tool schemas and history.
- Tools and APIs: search, CRM, payments, internal services, each with its own price or compute cost.
- Retrieval: embeddings for indexing plus vector queries on each relevant turn.
- Infrastructure: queue workers, sandboxes, storage and monitoring.
- Rework: retries, human review time and incident handling.
Most teams underestimate history growth and tool schemas, which inflate the prompt on every turn. Trimming both is usually the fastest saving available.
Pricing models and how to choose
Usage-based pricing aligns cost with work but makes forecasting hard, especially when agent traffic is bursty. Flat monthly pricing trades variable cost for predictability, which suits products with steady usage and teams that dislike per-token billing surprises. The honest answer is that no single model is cheapest for everyone: compare your measured cost per task against the plans on the live pricing page and re-check after major usage changes.
Plugsky self-serve plans are flat monthly with unlimited fair-use usage, and the free tier covers plugsky-micro and plugsky-lite with no card. A 14-day full-access trial lets you measure paid models on your real workload. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; batch endpoints, which typically reduce cost for bulk work, are coming soon.
Honest comparison
| Cost lever | Typical impact | Effort | Watch out for |
|---|---|---|---|
| Model routing per step | High | Low | Quality regressions on hard tasks |
| Turn and tool caps | High | Low | Cutting off tasks that needed one more step |
| Context and history trimming | Medium to high | Medium | Losing information the model needs |
| Prompt caching | Medium | Medium | Cache invalidation on prompt changes |
| Fewer, better tools | Medium | Medium | Removing tools users depend on |
| Flat monthly plans | Predictability | Low | Fair-use boundaries for bursty workloads |
Frequently asked questions
What is cost per task?
The total spend to complete one unit of work, including all model turns, tool calls, retrieval and retries. It is the metric that reflects real agent economics.
Why can a cheaper model cost more?
If it takes twice as many turns to finish, or fails and retries, the total token and tool cost can exceed a stronger model that completes the task quickly.
How do I reduce agent token usage?
Trim tool schemas and history, cache stable context, summarise old turns, cap tool result sizes and route routine steps to smaller models.
Does Plugsky charge per token?
Self-serve plans are flat monthly with unlimited fair-use usage rather than per-token billing. Current plan details are on the live pricing page.
Is there a free option?
Yes — plugsky-micro and plugsky-lite are free with no credit card, and a 14-day full-access trial covers paid models.
How do I forecast agent spend?
Instrument cost per task first, then multiply by expected task volume. For flat plans, compare that figure against plan tiers on the live pricing page.
Do batch endpoints reduce cost?
Batch processing is commonly cheaper because it trades latency for throughput; on Plugsky batch endpoints are coming soon, so plan bulk workloads accordingly.