Key facts
| Supervisor pattern | One orchestrator delegates tasks and merges results |
| Pipeline pattern | Agents run in sequence, each consuming the previous output |
| Debate pattern | Agents critique each other before a final answer is selected |
| Swarm pattern | Peers hand off directly based on declared capabilities |
| Shared state | A single typed state object beats agents messaging freely |
| Failure modes | Loops, deadlock, context loss, cost blow-up and error compounding |
| Model layer | 30+ models on one OpenAI-compatible key for per-agent routing |
| Governance | Scoped keys and audit logs are live for multi-agent traffic |
TL;DR
- Add agents only when one context cannot hold the work.
- Pick a known topology instead of inventing an interaction graph.
- Define handoff contracts and stop conditions before prompts.
- Evaluate intermediate steps; final-answer scoring hides failures.
- Cap cost per run and keep full traces of every handoff.
How it works, step by step
- Prove a single agent with tools fails or overloads before adding more agents.
- Choose a topology: supervisor, pipeline, debate or swarm.
- Define the shared state schema and what each agent may write to it.
- Specify handoff contracts: input shape, output shape and completion criteria.
- Set global limits: turn cap, token budget, retry ceiling and timeout.
- Evaluate step by step, with trajectory assertions for tool use and handoffs.
- Add human approval at the points where errors are expensive.
Try it yourself
Open the multi-agent workflow tool →
Choose a topology, not an improvisation
Most production multi-agent systems use one of four shapes. A supervisor owns the plan and delegates subtasks to specialists, then merges their results; this is the easiest to reason about and the most common. A pipeline runs agents in sequence, each transforming the previous output, which suits content, data and document flows. Debate has agents critique each other before a judge selects an answer, useful when correctness matters more than latency. A swarm lets peers hand off directly based on declared capabilities, which scales well but is the hardest to debug.
Pick one deliberately. Systems that emerge from ad-hoc messaging between agents accumulate ambiguity, and ambiguity in an agent system shows up as loops and duplicated work.
Shared state and handoff contracts
A single typed state object — with clear fields for plan, findings, artifacts and status — is easier to control than free-form message passing. Each agent reads what it needs and writes only to its own fields. Handoffs become contracts: the receiving agent knows the input shape, the output shape and what done means.
- Write scopes: one writer per field, so state never flickers between agents.
- Artifacts: large outputs go to storage with a reference in state, not into the prompt.
- Idempotency: a retried step must not duplicate side effects.
- Stop conditions: explicit completion checks per agent and for the whole run.
Failure modes, evaluation and cost
Multi-agent systems fail in predictable ways: loops where two agents bounce a task, deadlocks where each waits for the other, context loss at handoffs, error compounding where one wrong step poisons everything downstream, and cost blow-up when several agents each carry large context. Guard against all five with caps, idempotency and telemetry.
Evaluate step by step. Score tool selection, argument correctness and handoff accuracy, and keep full traces so a failed run can be replayed. Plugsky supports this pattern at the model layer: 30+ models on one OpenAI-compatible key with live function calling, streaming, JSON mode, embeddings and RAG, so each agent can run on the tier it needs. Scoped keys, RBAC, SSO/SCIM and audit logs cover access, though orchestration and tracing remain yours; assistants and batch endpoints are coming soon. Plans and the free plugsky-micro and plugsky-lite models are on the live pricing page.
Honest comparison
| Pattern | Best for | Main risk | Control point |
|---|---|---|---|
| Supervisor | Mixed tasks with one owner | Orchestrator context bloat | Delegation schema |
| Pipeline | Document and data flows | Error compounding | Stage validation gates |
| Debate | Correctness over latency | Cost and slower answers | Judge criteria |
| Swarm | Peer handoffs at scale | Hard debugging, loops | Capability registry |
| Single agent | Most tasks | Context limits | Tool design and turn cap |
Frequently asked questions
Do I need multiple agents?
Usually not. Start with one agent, good tools and retrieval. Split when a single context cannot hold the work, when roles need different tools or models, or when steps need separate evaluation.
Which topology should I start with?
Supervisor. It keeps one owner of the plan, makes failures traceable, and is simpler to cap and evaluate than peer handoffs.
How do agents share memory?
Prefer one typed state object with per-agent write scopes, plus external storage for large artifacts. Free-form conversation between agents is harder to control and audit.
How do I prevent loops?
Give every step a completion check, cap total turns and retries, and detect repeated identical states with a hash so the run stops instead of spinning.
How should I evaluate a multi-agent system?
Score each step: tool choice, argument validity, handoff correctness and state integrity. Add end-to-end task success as a separate metric. Final-answer scores alone hide broken intermediates.
Can each agent use a different model?
Yes, and it usually should. Route extraction and classification to small models, planning and synthesis to frontier models, all on one API key.
How do I control cost?
Cap tokens and turns per run, trim context at handoffs, avoid resending full histories, and track cost per completed task rather than per call.