Key facts
| Unit of evaluation | Task episode: goal, allowed tools, expected outcome and trajectory |
| Core metrics | Task success, tool-call correctness, argument validity, steps, latency, cost per task |
| Scoring | Deterministic assertions first; rubric LLM-as-judge for free-form output |
| Regression set | Frozen tasks re-run on every prompt, model or tool change |
| Online testing | Canary traffic with sampled human review before promotion |
| Model comparison | 30+ models behind one API makes A/B routing measurable |
| Free tier | Run evaluation harnesses on plugsky-micro and plugsky-lite at no cost |
| Roadmap | Managed evaluation and batch endpoints are coming soon; harnesses run now |
TL;DR
- Score trajectories: which tools were called, with what arguments, in how many steps.
- Use deterministic checks first and LLM judges only for open-ended quality.
- Freeze a regression set and run it on every prompt or model change.
- Track cost per completed task, not cost per token.
- Canary new versions on real traffic with sampled human review.
How it works, step by step
- Collect 30 to 100 real tasks with the tools the agent may use and the outcome you expect.
- Write assertions for hard requirements: correct tool, valid arguments, no forbidden actions.
- Add a rubric for qualities assertions cannot capture, and calibrate judges against human labels.
- Record per-task results: success, steps, tool errors, latency and cost.
- Run the suite after every prompt, model or tool change and block regressions in CI.
- Canary the winning configuration on a slice of live traffic with human review.
- Retire stale tasks and add new ones from real incidents and edge cases.
Try it yourself
Open the prompt diff and evaluator →
Evaluate the trajectory, not just the answer
Two agents can produce the same final answer while behaving very differently: one calls the right tool once, the other calls three wrong tools, ignores errors and guesses. If you only score the answer, both pass. Trajectory evaluation exposes that gap by checking the sequence of decisions — which tools were called, with which arguments, in what order and how many turns it took.
A practical episode record contains the goal, the allowed tool set, the messages, the tool calls with arguments and results, the final answer, and the metrics. That one structure supports assertions, judging, debugging and cost analysis.
Scoring layers
Layer deterministic checks first because they are cheap and unambiguous: did the agent call the lookup tool before answering, was the JSON valid, did it avoid any tool outside its allowlist, did it stop within the turn budget. Add rubric-based judging for qualities such as completeness or tone, calibrated against a small set of human-labelled examples so scores stay comparable over time.
- Hard requirements — pass or fail assertions that protect correctness and safety.
- Quality rubric — LLM judge with explicit criteria and examples.
- Efficiency — steps, tool calls and cost per successful task.
- Human review — mandatory for money-moving, legal or health workflows.
Keeping evaluation honest
Evaluation decays if it stops reflecting production. Rotate tasks from real traffic, keep a held-out set that never informs prompt tuning, and watch for judge drift by re-labelling a sample periodically. Track cost per completed task alongside quality: routing a task to a smaller model may keep success flat while cutting spend, but only measurement proves it.
On Plugsky, 30+ models sit behind one OpenAI-compatible endpoint, so comparing a free-tier model against a frontier model is a parameter change rather than a re-integration. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; batch and managed evaluation endpoints are coming soon. Plans, including the free plugsky-micro and plugsky-lite tier, are on the live pricing page.
Honest comparison
| Layer | What it catches | Cost to run | When to use |
|---|---|---|---|
| Deterministic assertions | Wrong tool, invalid JSON, forbidden action | Very low | Every run and every commit |
| Rubric LLM judge | Incomplete or poor-quality answers | Low to moderate | Before promotion and on samples |
| Trajectory scoring | Wasteful or fragile tool sequences | Low | Every regression run |
| Human review | Subtle failures and policy issues | High | High-stakes workflows and canaries |
| Online canary | Real-world regressions | Moderate | Before full rollout |
Frequently asked questions
How many test tasks do I need?
Start with 30 to 100 representative tasks covering the main workflows, edge cases and refusals. Grow the set from real incidents rather than inventing cases in bulk.
Can an LLM judge my agent?
Yes, for open-ended quality, with explicit rubrics and a handful of human-labelled calibration examples. Keep deterministic assertions for anything binary or safety-related.
What is the most useful single metric?
Task success rate alone hides waste. Pair success with cost and steps per completed task to catch agents that succeed by brute force.
How often should I re-run evaluations?
On every prompt, model or tool change in CI, plus a scheduled run to catch upstream model drift.
How do I test tool use without side effects?
Point tools at sandboxed or staging systems with idempotent operations, and assert on the calls the agent made rather than only the final state.
Should every model change trigger a rerun?
Yes. Different models call tools differently, so the same prompt can regress even when answers look similar.
Is there a free way to run evals?
Yes. plugsky-micro and plugsky-lite are free with no card, so harnesses and smoke tests can run continuously. A 14-day full-access trial covers paid models.