Agents

How much do AI agents cost to run?

Agent cost is the sum of six buckets: model tokens, tool and API calls, sandbox or browser compute, storage and memory, observability and evaluation, and human review. The dominant variable is usually tokens, because agents re-read context on every step. Control it with step caps, model routing, context trimming, caching and cost per completed task as the metric — not cost per token.

Key facts

Model tokensInput, output and reasoning tokens across every step of a run
Tool callsSearch APIs, enrichment providers, databases and internal services
ComputeSandbox VMs, browser sessions and queue workers for long runs
MemoryVector store, object storage and trace retention costs
Waste sourcesRetries, unbounded loops, oversized context and wrong model tier
Pricing modelPlugsky self-serve plans are flat monthly with unlimited fair-use usage
Free tierplugsky-micro and plugsky-lite free, no card required
MetricTrack cost per completed task, not cost per token

TL;DR

  • Tokens dominate agent cost because context is re-sent every step.
  • Tool calls and sandbox compute are the second and third buckets.
  • Most overspend comes from retries, loops and wrong model tiers.
  • Cap steps, trim context, cache aggressively and route by task difficulty.
  • The only honest metric is cost per completed task.

How it works, step by step

  1. Instrument every step: model, tokens in and out, tool calls, latency and outcome.
  2. Compute cost per completed task for each workflow and rank them.
  3. Set hard limits: maximum steps, wall-clock timeout and token budget per run.
  4. Trim context between steps instead of resending the full transcript.
  5. Route simple steps to small models and reserve frontier models for reasoning.
  6. Cache deterministic tool results and reuse retrieval across runs.
  7. Review the ranking monthly and delete or fix the worst-performing workflow.
1Instrument everystep: model, tokensin and out, tool2Compute cost percompleted task foreach workflow and3Set hard limits:maximum steps,wall-clock timeout4Trim contextbetween stepsinstead of5Route simple stepsto small models andreserve frontier6Cache deterministictool results andreuse retrieval

Try it yourself

Open the agent run cost calculator →

The six cost buckets

Model tokens come first and usually dominate. An agent that takes fifteen steps re-sends its context fifteen times, so input tokens scale with steps and with how much history you carry. Output and reasoning tokens add on top, and reasoning-heavy models can spend many tokens before producing a visible action.

Then the supporting costs: tool calls to search, enrichment or internal APIs, many of which are metered; sandbox compute for code execution, browser sessions and background workers; storage for vectors, artifacts and retained traces; observability and evaluation tooling; and human review time, which is a real cost even when it does not appear on a cloud bill.

Where the money leaks

Overspend is rarely caused by list prices. It comes from agent behaviour: loops that never terminate, retries that repeat an identical failing call, full conversation history pasted into every step, retrieval that returns twenty documents when three would do, and a frontier model doing classification work a small model handles just as well.

  • Unbounded loops: always set a step cap and a budget per run.
  • Retry storms: exponential backoff with a retry ceiling, and never retry a deterministic failure.
  • Context bloat: summarise or trim between steps; keep the working set small.
  • Wrong tier: route extraction, formatting and routing decisions to cheap models.
  • Failed runs: count them; they are pure cost with no output.

How to model and control cost

Build a simple model: average steps per task, average tokens per step, price per token for the model tier, plus tool and compute costs. That gives an expected cost per task that you can compare against its business value. Then attack the biggest multiplier first — usually steps or context size.

Platform pricing shapes the rest. Plugsky self-serve plans are flat monthly with unlimited fair-use usage rather than per-token billing, which removes token price as a variable on the platform side and makes forecasting a function of your own usage. With 30+ models behind one OpenAI-compatible key, each step can also be routed to the cheapest tier that completes it reliably. The free tier covers plugsky-micro and plugsky-lite, and a 14-day full-access trial exists; current plans are on the live pricing page. Use the cost calculator to estimate per-run spend before you scale a workflow.

Honest comparison

Cost driverWhat pushes it upControlMetric to watch
Model tokensMany steps, long context, big modelStep caps, trimming, routing, cachingTokens per completed task
Tool callsChatty agents, redundant lookupsCache results, batch calls, narrow toolsTool calls per task
Sandbox computeLong browser and code sessionsTimeouts, snapshots, right-size VMsCompute minutes per run
StorageUnpruned vectors and tracesRetention policy, summarise old stateCost per stored task
Human reviewLow-confidence automationConfidence thresholds, better toolingMinutes reviewed per task

Frequently asked questions

What is the biggest cost in an AI agent?

Model tokens, because context is re-sent on every step. Steps and context size multiply, so trimming either one usually cuts the bill fastest.

How do I calculate cost per task?

Sum tokens across all steps in a run, multiply by model price, then add tool, compute and storage costs for that run, and divide by the number of tasks completed successfully.

Is per-token or flat pricing better for agents?

It depends on variability. Flat platform plans make forecasting simple; per-token billing tracks usage precisely. Plugsky self-serve plans are flat monthly with fair-use usage.

How do I stop runaway loops?

Set a maximum number of steps, a wall-clock timeout and a token budget per run, and require the agent to satisfy a completion check before it can continue.

Do failed runs cost money?

Yes, and they are easy to ignore. Track failure rate per workflow and treat each failure as cost with no output; reducing failures is often the cheapest optimisation.

Should I always use the cheapest model?

No. Use the cheapest model that reliably completes each step. Cheap models that fail and retry cost more than a mid-tier model that succeeds first time.

Where can I estimate costs before building?

Use the agent run cost calculator to model steps, tokens and tool calls, then validate the estimate against instrumentation from a real pilot run.