Key facts
| Failure layers | Model output, tool execution, network transport, workflow state |
| Retry policy | Bounded retries with exponential backoff and jitter for transient errors |
| Idempotency | Idempotency keys prevent duplicate sends, payments and writes on retry |
| Checkpointing | Persist run state after each step so a failed run can resume |
| Timeouts | Per-tool timeout plus a total turn and wall-clock budget |
| Circuit breaker | Stop calling a failing dependency and fail over to a fallback path |
| Dead letter | Park unrecoverable runs for human review with full trace context |
| Live API | Streaming, JSON mode and function calling are live on the agent API |
TL;DR
- Classify failures before retrying: only transient errors deserve an automatic retry.
- Use exponential backoff with jitter and a hard attempt cap.
- Make side effects idempotent so retries cannot double-charge or double-send.
- Checkpoint each step so a long run resumes instead of starting over.
- Route unrecoverable runs to a dead-letter queue with a human owner.
How it works, step by step
- Define an error taxonomy: transient, permanent, policy refusal and malformed model output.
- Set retry budgets per error class with exponential backoff, jitter and a maximum attempt count.
- Require idempotency keys on every tool that writes, sends or pays.
- Checkpoint state after each successful step so runs can resume.
- Add per-tool timeouts, circuit breakers and a global wall-clock budget.
- Send exhausted runs to a dead-letter queue with the trace attached and alert a human owner.
- Replay failed runs from fixtures and add each incident to the evaluation set.
Try it yourself
Open the OpenAI error decoder →
Four layers where agents fail
Model failures look like malformed tool arguments, refusals or loops that never converge. Tool failures look like timeouts, 5xx responses and rate limits. Network failures are dropped connections and slow DNS. Workflow failures are the worst: the process dies halfway through a multi-step task with side effects already applied. Each layer needs a different response, so classify before you retry.
- Retry: transient network errors, rate limits and 5xx responses.
- Repair: malformed model output — re-prompt with the validation error attached.
- Escalate: policy refusals and permission denials — stop and tell the user.
- Recover: process crashes — resume from the last checkpoint.
Retry, backoff and idempotency
Naive retries multiply load exactly when a dependency is struggling. Use exponential backoff with jitter, cap the attempts, and only retry errors that are safe to repeat. The moment a tool has a side effect — sending an email, charging a card, creating an order — the retry story changes: the operation must be idempotent, keyed by a value the agent generates once per logical action and reuses on every attempt.
Pair retries with a circuit breaker per dependency. After a threshold of consecutive failures, stop calling the broken tool, return a clear message to the model and let the run take a fallback path or fail gracefully with a useful explanation instead of burning the turn budget on timeouts.
Checkpoints and the dead-letter path
Long agent runs should be durable state machines, not a single in-memory loop. Persist the message history, completed steps and tool results after each step so a crash resumes rather than restarts. If the agent moved money or deleted data before crashing, the resume path must know what already applied — that is the idempotency ledger's job.
When retries are exhausted or the agent loops without progress, stop and hand off. A dead-letter queue with the full trace, the failing step and the last error gives an operator what they need to fix, replay or finish the task manually. Feed every incident back into the evaluation set so the same class of failure is caught before release. Teams running this on Plugsky can lean on 30+ models behind one endpoint to fail over between tiers, with plans on the live pricing page.
Honest comparison
| Failure | Wrong response | Right response | Guardrail |
|---|---|---|---|
| Tool timeout | Immediate unbounded retry | Backoff retry, then fallback | Per-tool timeout and attempt cap |
| Malformed arguments | Send raw output back unchanged | Re-prompt with validation error | Schema validation before execution |
| Duplicate side effect | Retry the write | Idempotency key on the action | Ledger of applied actions |
| Process crash mid-run | Restart from step one | Resume from checkpoint | Durable step state |
| Repeated policy refusal | Loop and rephrase forever | Stop and escalate to user | Turn budget and refusal handling |
Frequently asked questions
Which errors should an agent retry?
Transient ones: dropped connections, rate limits and 5xx responses. Do not automatically retry validation errors, permission denials or policy refusals.
What backoff should I use?
Exponential backoff with random jitter and a hard attempt cap per tool call, plus a total wall-clock budget for the whole run.
How do idempotency keys work?
Generate one key per logical action, send it with the tool call, and have the tool return the original result if the key was already processed. Retries then cannot double-apply.
How do I resume a crashed run?
Persist message history, completed steps and results after each step, then restart the loop from the last checkpoint with the same idempotency keys.
When should the agent give up?
When retries are exhausted, the turn budget is spent, or the agent repeats the same failing action. Park the run for review rather than looping.
Does retrying increase cost?
Yes. Retries consume tokens and tool calls, which is why bounded budgets and circuit breakers matter for per-task cost control.
Can Plugsky help with failover?
Yes. The OpenAI-compatible API gives access to 30+ models, so you can configure fallback routing between tiers without changing your agent code.