Agents

How do you evaluate an AI agent properly?

Evaluate agents on trajectories, not only final answers. Build a labelled set of tasks with expected tool calls and outcomes, then score task success, tool-call correctness, argument validity, steps taken, latency and cost per task. Run deterministic checks first, use rubric-based LLM judging for open answers, add human review for high-stakes runs, and re-run the set on every prompt or model change.

Key facts

Unit of evaluationTask episode: goal, allowed tools, expected outcome and trajectory
Core metricsTask success, tool-call correctness, argument validity, steps, latency, cost per task
ScoringDeterministic assertions first; rubric LLM-as-judge for free-form output
Regression setFrozen tasks re-run on every prompt, model or tool change
Online testingCanary traffic with sampled human review before promotion
Model comparison30+ models behind one API makes A/B routing measurable
Free tierRun evaluation harnesses on plugsky-micro and plugsky-lite at no cost
RoadmapManaged evaluation and batch endpoints are coming soon; harnesses run now

TL;DR

  • Score trajectories: which tools were called, with what arguments, in how many steps.
  • Use deterministic checks first and LLM judges only for open-ended quality.
  • Freeze a regression set and run it on every prompt or model change.
  • Track cost per completed task, not cost per token.
  • Canary new versions on real traffic with sampled human review.

How it works, step by step

  1. Collect 30 to 100 real tasks with the tools the agent may use and the outcome you expect.
  2. Write assertions for hard requirements: correct tool, valid arguments, no forbidden actions.
  3. Add a rubric for qualities assertions cannot capture, and calibrate judges against human labels.
  4. Record per-task results: success, steps, tool errors, latency and cost.
  5. Run the suite after every prompt, model or tool change and block regressions in CI.
  6. Canary the winning configuration on a slice of live traffic with human review.
  7. Retire stale tasks and add new ones from real incidents and edge cases.
1Collect 30 to 100real tasks with thetools the agent may2Write assertionsfor hardrequirements:3Add a rubric forqualitiesassertions cannot4Record per-taskresults: success,steps, tool errors,5Run the suite afterevery prompt, modelor tool change and6Canary the winningconfiguration on aslice of live

Try it yourself

Open the prompt diff and evaluator →

Evaluate the trajectory, not just the answer

Two agents can produce the same final answer while behaving very differently: one calls the right tool once, the other calls three wrong tools, ignores errors and guesses. If you only score the answer, both pass. Trajectory evaluation exposes that gap by checking the sequence of decisions — which tools were called, with which arguments, in what order and how many turns it took.

A practical episode record contains the goal, the allowed tool set, the messages, the tool calls with arguments and results, the final answer, and the metrics. That one structure supports assertions, judging, debugging and cost analysis.

Scoring layers

Layer deterministic checks first because they are cheap and unambiguous: did the agent call the lookup tool before answering, was the JSON valid, did it avoid any tool outside its allowlist, did it stop within the turn budget. Add rubric-based judging for qualities such as completeness or tone, calibrated against a small set of human-labelled examples so scores stay comparable over time.

  • Hard requirements — pass or fail assertions that protect correctness and safety.
  • Quality rubric — LLM judge with explicit criteria and examples.
  • Efficiency — steps, tool calls and cost per successful task.
  • Human review — mandatory for money-moving, legal or health workflows.

Keeping evaluation honest

Evaluation decays if it stops reflecting production. Rotate tasks from real traffic, keep a held-out set that never informs prompt tuning, and watch for judge drift by re-labelling a sample periodically. Track cost per completed task alongside quality: routing a task to a smaller model may keep success flat while cutting spend, but only measurement proves it.

On Plugsky, 30+ models sit behind one OpenAI-compatible endpoint, so comparing a free-tier model against a frontier model is a parameter change rather than a re-integration. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; batch and managed evaluation endpoints are coming soon. Plans, including the free plugsky-micro and plugsky-lite tier, are on the live pricing page.

Honest comparison

LayerWhat it catchesCost to runWhen to use
Deterministic assertionsWrong tool, invalid JSON, forbidden actionVery lowEvery run and every commit
Rubric LLM judgeIncomplete or poor-quality answersLow to moderateBefore promotion and on samples
Trajectory scoringWasteful or fragile tool sequencesLowEvery regression run
Human reviewSubtle failures and policy issuesHighHigh-stakes workflows and canaries
Online canaryReal-world regressionsModerateBefore full rollout

Frequently asked questions

How many test tasks do I need?

Start with 30 to 100 representative tasks covering the main workflows, edge cases and refusals. Grow the set from real incidents rather than inventing cases in bulk.

Can an LLM judge my agent?

Yes, for open-ended quality, with explicit rubrics and a handful of human-labelled calibration examples. Keep deterministic assertions for anything binary or safety-related.

What is the most useful single metric?

Task success rate alone hides waste. Pair success with cost and steps per completed task to catch agents that succeed by brute force.

How often should I re-run evaluations?

On every prompt, model or tool change in CI, plus a scheduled run to catch upstream model drift.

How do I test tool use without side effects?

Point tools at sandboxed or staging systems with idempotent operations, and assert on the calls the agent made rather than only the final state.

Should every model change trigger a rerun?

Yes. Different models call tools differently, so the same prompt can regress even when answers look similar.

Is there a free way to run evals?

Yes. plugsky-micro and plugsky-lite are free with no card, so harnesses and smoke tests can run continuously. A 14-day full-access trial covers paid models.