Key facts
| Model tokens | Input, output and reasoning tokens across every step of a run |
| Tool calls | Search APIs, enrichment providers, databases and internal services |
| Compute | Sandbox VMs, browser sessions and queue workers for long runs |
| Memory | Vector store, object storage and trace retention costs |
| Waste sources | Retries, unbounded loops, oversized context and wrong model tier |
| Pricing model | Plugsky self-serve plans are flat monthly with unlimited fair-use usage |
| Free tier | plugsky-micro and plugsky-lite free, no card required |
| Metric | Track cost per completed task, not cost per token |
TL;DR
- Tokens dominate agent cost because context is re-sent every step.
- Tool calls and sandbox compute are the second and third buckets.
- Most overspend comes from retries, loops and wrong model tiers.
- Cap steps, trim context, cache aggressively and route by task difficulty.
- The only honest metric is cost per completed task.
How it works, step by step
- Instrument every step: model, tokens in and out, tool calls, latency and outcome.
- Compute cost per completed task for each workflow and rank them.
- Set hard limits: maximum steps, wall-clock timeout and token budget per run.
- Trim context between steps instead of resending the full transcript.
- Route simple steps to small models and reserve frontier models for reasoning.
- Cache deterministic tool results and reuse retrieval across runs.
- Review the ranking monthly and delete or fix the worst-performing workflow.
Try it yourself
Open the agent run cost calculator →
The six cost buckets
Model tokens come first and usually dominate. An agent that takes fifteen steps re-sends its context fifteen times, so input tokens scale with steps and with how much history you carry. Output and reasoning tokens add on top, and reasoning-heavy models can spend many tokens before producing a visible action.
Then the supporting costs: tool calls to search, enrichment or internal APIs, many of which are metered; sandbox compute for code execution, browser sessions and background workers; storage for vectors, artifacts and retained traces; observability and evaluation tooling; and human review time, which is a real cost even when it does not appear on a cloud bill.
Where the money leaks
Overspend is rarely caused by list prices. It comes from agent behaviour: loops that never terminate, retries that repeat an identical failing call, full conversation history pasted into every step, retrieval that returns twenty documents when three would do, and a frontier model doing classification work a small model handles just as well.
- Unbounded loops: always set a step cap and a budget per run.
- Retry storms: exponential backoff with a retry ceiling, and never retry a deterministic failure.
- Context bloat: summarise or trim between steps; keep the working set small.
- Wrong tier: route extraction, formatting and routing decisions to cheap models.
- Failed runs: count them; they are pure cost with no output.
How to model and control cost
Build a simple model: average steps per task, average tokens per step, price per token for the model tier, plus tool and compute costs. That gives an expected cost per task that you can compare against its business value. Then attack the biggest multiplier first — usually steps or context size.
Platform pricing shapes the rest. Plugsky self-serve plans are flat monthly with unlimited fair-use usage rather than per-token billing, which removes token price as a variable on the platform side and makes forecasting a function of your own usage. With 30+ models behind one OpenAI-compatible key, each step can also be routed to the cheapest tier that completes it reliably. The free tier covers plugsky-micro and plugsky-lite, and a 14-day full-access trial exists; current plans are on the live pricing page. Use the cost calculator to estimate per-run spend before you scale a workflow.
Honest comparison
| Cost driver | What pushes it up | Control | Metric to watch |
|---|---|---|---|
| Model tokens | Many steps, long context, big model | Step caps, trimming, routing, caching | Tokens per completed task |
| Tool calls | Chatty agents, redundant lookups | Cache results, batch calls, narrow tools | Tool calls per task |
| Sandbox compute | Long browser and code sessions | Timeouts, snapshots, right-size VMs | Compute minutes per run |
| Storage | Unpruned vectors and traces | Retention policy, summarise old state | Cost per stored task |
| Human review | Low-confidence automation | Confidence thresholds, better tooling | Minutes reviewed per task |
Frequently asked questions
What is the biggest cost in an AI agent?
Model tokens, because context is re-sent on every step. Steps and context size multiply, so trimming either one usually cuts the bill fastest.
How do I calculate cost per task?
Sum tokens across all steps in a run, multiply by model price, then add tool, compute and storage costs for that run, and divide by the number of tasks completed successfully.
Is per-token or flat pricing better for agents?
It depends on variability. Flat platform plans make forecasting simple; per-token billing tracks usage precisely. Plugsky self-serve plans are flat monthly with fair-use usage.
How do I stop runaway loops?
Set a maximum number of steps, a wall-clock timeout and a token budget per run, and require the agent to satisfy a completion check before it can continue.
Do failed runs cost money?
Yes, and they are easy to ignore. Track failure rate per workflow and treat each failure as cost with no output; reducing failures is often the cheapest optimisation.
Should I always use the cheapest model?
No. Use the cheapest model that reliably completes each step. Cheap models that fail and retry cost more than a mid-tier model that succeeds first time.
Where can I estimate costs before building?
Use the agent run cost calculator to model steps, tokens and tool calls, then validate the estimate against instrumentation from a real pilot run.