Key facts
| Pattern one | Supervisor routes work to specialised worker agents |
| Pattern two | Deterministic workflow with agent steps for open-ended tasks |
| State | Shared run state with checkpoints so long runs can resume |
| Handoffs | Pass a structured brief, not the entire transcript |
| Budgets | Turn, token, time and cost caps per run and per branch |
| Tracing | One trace ID across router, agents, tools and services |
| Model routing | 30+ models behind one API key, chosen per task difficulty |
| Status | Function calling, streaming, JSON mode and agents are live |
TL;DR
- Prefer a deterministic workflow over an autonomous loop wherever steps are known.
- Use a supervisor only when work genuinely branches across specialised agents.
- Hand off structured briefs, not raw transcripts, to control context and cost.
- Budget every branch with turn, time and cost caps, and checkpoint state.
- Trace the whole run under one ID so failures are debuggable end to end.
How it works, step by step
- Map the task as a flow and mark which steps are fixed and which need judgement.
- Implement fixed steps as ordinary code and use model calls only where the path is unknown.
- Add a router or supervisor when you need to select between distinct workflows.
- Define handoff payloads as schemas so sub-agents receive only what they need.
- Persist run state after each step and make side effects idempotent.
- Set turn, token, time and cost budgets per run and per branch.
- Emit one trace ID across router, agents, tools and downstream services, and evaluate the flow end to end.
Try it yourself
Open the agent workflow designer →
Start with the simplest structure that works
Autonomy is a spectrum, not a binary. A single model call with a good prompt solves more tasks than teams expect. A prompt chain with two or three fixed steps solves more. A tool-calling loop handles genuinely open-ended work. Multi-agent systems are justified only when different steps need different tools, prompts or models, and when no single agent can hold the context.
Each level adds failure surface: more prompts to maintain, more state to persist, more cost per task and harder debugging. Choose the lowest level that meets the requirement, and treat every promotion to a more autonomous architecture as a decision that must earn its complexity in evaluation results.
Supervisors, workers and handoffs
The supervisor pattern puts one routing agent in front of specialised workers: a billing agent with billing tools, a research agent with search tools, a scheduling agent with calendar access. The supervisor classifies the request and hands off. The workers stay narrow, which keeps their prompts stable and their tool permissions small.
- Route on outcome, not on keywords alone; let the model classify with a constrained label set.
- Hand off a brief: goal, constraints, relevant facts and the expected output format.
- Cap depth: limit how many handoffs a request can trigger to prevent loops.
- Return control: workers report back to the supervisor rather than talking to each other directly.
Budgets, state and tracing
Orchestration without budgets becomes a cost incident. Give every run a maximum number of turns, a token ceiling, a wall-clock deadline and, where possible, a monetary estimate. Checkpoint state after each step so a failure resumes instead of restarting, and make every side effect idempotent so replays are safe.
Observability is what makes this operable: one trace ID that links the routing decision, each agent's turns, every tool call and the final outcome. On Plugsky you can route each step to the right tier from 30+ models on one OpenAI-compatible key, keeping cheap models on classification and frontier models on planning. Chat, streaming, JSON mode, function calling, embeddings, RAG and agents are live; managed assistants and batch endpoints are coming soon. Plans are on the live pricing page.
Honest comparison
| Approach | Best for | Main failure mode | Use when |
|---|---|---|---|
| Single call | Classification, extraction, short answers | Cannot act on results | No tool use needed |
| Prompt chain | Known multi-step workflows | Brittle to input variation | Steps are fixed |
| Tool loop | Open-ended tasks with clear tools | Loops or wasted turns | Path must be discovered |
| Supervisor plus workers | Distinct domains and tool sets | Routing errors and cost | One agent cannot hold context |
| Deterministic workflow with agent steps | Regulated or repeatable processes | Over-constraining the model | Audit and repeatability required |
Frequently asked questions
Do I need multiple agents?
Usually not. Start with one agent and a small tool set. Add workers only when steps need different tools, prompts or models, or when context no longer fits one run.
What is the supervisor pattern?
A routing agent receives the request, classifies it and hands off to a specialised worker agent. Workers report back, and the supervisor owns the final response.
How do I stop runaway agents?
Set turn, token, time and cost budgets per run and per branch, cap handoff depth, and stop when the agent repeats the same failing action.
Should agents share one conversation history?
No. Pass a structured brief to the next agent so its context stays small, relevant and cheaper to process.
How do I debug a multi-step run?
Use one trace ID across every model call, tool call and service, then replay the trace against your evaluation set after each fix.
Can I use LangChain or LangGraph with Plugsky?
Yes. The API is OpenAI-compatible, so frameworks work after a base URL change, and you keep orchestration in code you control.
Which models should each step use?
Route routine classification and rewriting to plugsky-micro or plugsky-lite, and reserve frontier models for planning, with 30+ models available on one API key.