Agents

How is agentic AI priced?

Agent cost is not token price alone. The real unit is cost per completed task: tokens per turn multiplied by turns, plus tool and API calls, retrieval, infrastructure and retries. Reduce it with smaller models on routine steps, caching, tighter context, turn caps and fewer tools. Plugsky self-serve plans are flat monthly with unlimited fair-use usage — see the live pricing page.

Key facts

Cost unitCost per completed task, not cost per token
Main driversTurns multiplied by tokens, plus tools, retrieval and retries
Reduction leversModel routing, caching, context trimming, turn caps
Self-serve billingFlat monthly plans with unlimited fair-use usage
Free tierplugsky-micro and plugsky-lite free, no card required
Trial14-day full-access trial for paid models
VisibilityUsage dashboard shows tokens, spend and model mix
Current amountsLive prices are published on the Plugsky pricing page

TL;DR

  • Budget per completed task, because a cheap model that loops can cost more than a strong one that finishes.
  • Turn count is the biggest multiplier: cap it and count tools per run.
  • Route routine steps to small models and reserve frontier models for planning.
  • Cache stable context and trim tool results to cut tokens on every turn.
  • Flat monthly plans remove per-token forecasting; check the live pricing page.

How it works, step by step

  1. Instrument the agent to record tokens, turns, tool calls and retries per completed task.
  2. Calculate a baseline cost per task for each workflow before optimising anything.
  3. Route classification and rewriting to the cheapest model that passes your evaluations.
  4. Cache stable context such as policies and schemas, and trim verbose tool results.
  5. Set turn and tool-call caps so a confused run stops early instead of billing on.
  6. Re-run the task set after each change to confirm quality held while cost fell.
  7. Compare usage against flat plan tiers on the live pricing page and move when the math favours it.
1Instrument theagent to recordtokens, turns, tool2Calculate abaseline cost pertask for each3Routeclassification andrewriting to the4Cache stablecontext such aspolicies and5Set turn andtool-call caps so aconfused run stops6Re-run the task setafter each changeto confirm quality

Try it yourself

Open the agent run cost calculator →

Why agent pricing is different

A single chat request costs one prompt and one completion. An agent might make ten model calls, run eight tools, retrieve documents twice, fail once and retry. Every one of those steps consumes something: tokens, API calls, compute, database reads or third-party fees. Pricing an agent per token hides that structure; pricing it per completed task exposes it.

That reframing changes optimisation too. A slightly more expensive model that finishes in three turns can beat a cheaper model that wanders for twelve. Measure outcomes, not unit prices.

The cost stack

  • Model tokens: turns multiplied by input and output tokens, including tool schemas and history.
  • Tools and APIs: search, CRM, payments, internal services, each with its own price or compute cost.
  • Retrieval: embeddings for indexing plus vector queries on each relevant turn.
  • Infrastructure: queue workers, sandboxes, storage and monitoring.
  • Rework: retries, human review time and incident handling.

Most teams underestimate history growth and tool schemas, which inflate the prompt on every turn. Trimming both is usually the fastest saving available.

Pricing models and how to choose

Usage-based pricing aligns cost with work but makes forecasting hard, especially when agent traffic is bursty. Flat monthly pricing trades variable cost for predictability, which suits products with steady usage and teams that dislike per-token billing surprises. The honest answer is that no single model is cheapest for everyone: compare your measured cost per task against the plans on the live pricing page and re-check after major usage changes.

Plugsky self-serve plans are flat monthly with unlimited fair-use usage, and the free tier covers plugsky-micro and plugsky-lite with no card. A 14-day full-access trial lets you measure paid models on your real workload. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; batch endpoints, which typically reduce cost for bulk work, are coming soon.

Honest comparison

Cost leverTypical impactEffortWatch out for
Model routing per stepHighLowQuality regressions on hard tasks
Turn and tool capsHighLowCutting off tasks that needed one more step
Context and history trimmingMedium to highMediumLosing information the model needs
Prompt cachingMediumMediumCache invalidation on prompt changes
Fewer, better toolsMediumMediumRemoving tools users depend on
Flat monthly plansPredictabilityLowFair-use boundaries for bursty workloads

Frequently asked questions

What is cost per task?

The total spend to complete one unit of work, including all model turns, tool calls, retrieval and retries. It is the metric that reflects real agent economics.

Why can a cheaper model cost more?

If it takes twice as many turns to finish, or fails and retries, the total token and tool cost can exceed a stronger model that completes the task quickly.

How do I reduce agent token usage?

Trim tool schemas and history, cache stable context, summarise old turns, cap tool result sizes and route routine steps to smaller models.

Does Plugsky charge per token?

Self-serve plans are flat monthly with unlimited fair-use usage rather than per-token billing. Current plan details are on the live pricing page.

Is there a free option?

Yes — plugsky-micro and plugsky-lite are free with no credit card, and a 14-day full-access trial covers paid models.

How do I forecast agent spend?

Instrument cost per task first, then multiply by expected task volume. For flat plans, compare that figure against plan tiers on the live pricing page.

Do batch endpoints reduce cost?

Batch processing is commonly cheaper because it trades latency for throughput; on Plugsky batch endpoints are coming soon, so plan bulk workloads accordingly.