Agents

How should AI agents handle failures and retries?

Agents fail at four layers: model, tool, network and workflow state. Design for each with bounded retries and exponential backoff for transient errors, idempotency keys so retried side effects happen once, checkpoints so a run resumes instead of restarting, per-tool timeouts and circuit breakers, and a dead-letter queue that parks unrecoverable runs for human review.

Key facts

Failure layersModel output, tool execution, network transport, workflow state
Retry policyBounded retries with exponential backoff and jitter for transient errors
IdempotencyIdempotency keys prevent duplicate sends, payments and writes on retry
CheckpointingPersist run state after each step so a failed run can resume
TimeoutsPer-tool timeout plus a total turn and wall-clock budget
Circuit breakerStop calling a failing dependency and fail over to a fallback path
Dead letterPark unrecoverable runs for human review with full trace context
Live APIStreaming, JSON mode and function calling are live on the agent API

TL;DR

  • Classify failures before retrying: only transient errors deserve an automatic retry.
  • Use exponential backoff with jitter and a hard attempt cap.
  • Make side effects idempotent so retries cannot double-charge or double-send.
  • Checkpoint each step so a long run resumes instead of starting over.
  • Route unrecoverable runs to a dead-letter queue with a human owner.

How it works, step by step

  1. Define an error taxonomy: transient, permanent, policy refusal and malformed model output.
  2. Set retry budgets per error class with exponential backoff, jitter and a maximum attempt count.
  3. Require idempotency keys on every tool that writes, sends or pays.
  4. Checkpoint state after each successful step so runs can resume.
  5. Add per-tool timeouts, circuit breakers and a global wall-clock budget.
  6. Send exhausted runs to a dead-letter queue with the trace attached and alert a human owner.
  7. Replay failed runs from fixtures and add each incident to the evaluation set.
1Define an errortaxonomy:transient,2Set retry budgetsper error classwith exponential3Require idempotencykeys on every toolthat writes, sends4Checkpoint stateafter eachsuccessful step so5Add per-tooltimeouts, circuitbreakers and a6Send exhausted runsto a dead-letterqueue with the

Try it yourself

Open the OpenAI error decoder →

Four layers where agents fail

Model failures look like malformed tool arguments, refusals or loops that never converge. Tool failures look like timeouts, 5xx responses and rate limits. Network failures are dropped connections and slow DNS. Workflow failures are the worst: the process dies halfway through a multi-step task with side effects already applied. Each layer needs a different response, so classify before you retry.

  • Retry: transient network errors, rate limits and 5xx responses.
  • Repair: malformed model output — re-prompt with the validation error attached.
  • Escalate: policy refusals and permission denials — stop and tell the user.
  • Recover: process crashes — resume from the last checkpoint.

Retry, backoff and idempotency

Naive retries multiply load exactly when a dependency is struggling. Use exponential backoff with jitter, cap the attempts, and only retry errors that are safe to repeat. The moment a tool has a side effect — sending an email, charging a card, creating an order — the retry story changes: the operation must be idempotent, keyed by a value the agent generates once per logical action and reuses on every attempt.

Pair retries with a circuit breaker per dependency. After a threshold of consecutive failures, stop calling the broken tool, return a clear message to the model and let the run take a fallback path or fail gracefully with a useful explanation instead of burning the turn budget on timeouts.

Checkpoints and the dead-letter path

Long agent runs should be durable state machines, not a single in-memory loop. Persist the message history, completed steps and tool results after each step so a crash resumes rather than restarts. If the agent moved money or deleted data before crashing, the resume path must know what already applied — that is the idempotency ledger's job.

When retries are exhausted or the agent loops without progress, stop and hand off. A dead-letter queue with the full trace, the failing step and the last error gives an operator what they need to fix, replay or finish the task manually. Feed every incident back into the evaluation set so the same class of failure is caught before release. Teams running this on Plugsky can lean on 30+ models behind one endpoint to fail over between tiers, with plans on the live pricing page.

Honest comparison

FailureWrong responseRight responseGuardrail
Tool timeoutImmediate unbounded retryBackoff retry, then fallbackPer-tool timeout and attempt cap
Malformed argumentsSend raw output back unchangedRe-prompt with validation errorSchema validation before execution
Duplicate side effectRetry the writeIdempotency key on the actionLedger of applied actions
Process crash mid-runRestart from step oneResume from checkpointDurable step state
Repeated policy refusalLoop and rephrase foreverStop and escalate to userTurn budget and refusal handling

Frequently asked questions

Which errors should an agent retry?

Transient ones: dropped connections, rate limits and 5xx responses. Do not automatically retry validation errors, permission denials or policy refusals.

What backoff should I use?

Exponential backoff with random jitter and a hard attempt cap per tool call, plus a total wall-clock budget for the whole run.

How do idempotency keys work?

Generate one key per logical action, send it with the tool call, and have the tool return the original result if the key was already processed. Retries then cannot double-apply.

How do I resume a crashed run?

Persist message history, completed steps and results after each step, then restart the loop from the last checkpoint with the same idempotency keys.

When should the agent give up?

When retries are exhausted, the turn budget is spent, or the agent repeats the same failing action. Park the run for review rather than looping.

Does retrying increase cost?

Yes. Retries consume tokens and tool calls, which is why bounded budgets and circuit breakers matter for per-task cost control.

Can Plugsky help with failover?

Yes. The OpenAI-compatible API gives access to 30+ models, so you can configure fallback routing between tiers without changing your agent code.