Back to insights
Field Notes·9 min read

Naive Retries Are Quietly Wrecking Your Production Agent

Naive Retries Are Quietly Wrecking Your Production Agent

Most agent retry logic is a try/except with a loop and a sleep, copied from a Stack Overflow answer about flaky HTTP. It works fine against transient network blips. It falls apart the moment the thing you're retrying already changed state on the first attempt — and in an agent, half the interesting tool calls do exactly that. We keep finding the same pattern in audits: a retry loop that turned one failed action into three completed ones, and nobody noticed until the second invoice went out.

By Daniel Usvyat · Principal, USQRD

Share

The Retry That Charged the Card Twice

Here's the mechanism, because it's more mundane than people expect. An agent calls a `create_charge` tool. The tool hits the payment API, the charge succeeds, and then the response times out on the way back — the connection drops after the mutation but before the agent sees the 200. The agent's retry wrapper sees a timeout, assumes failure, and calls `create_charge` again. Now there are two charges and one very confused customer.

The API did nothing wrong. The payment succeeded both times. The bug is that the retry code treated a non-idempotent operation as if it were safe to repeat. Timeouts are ambiguous by nature — a timeout tells you the response didn't arrive. It doesn't tell you whether the work happened. Naive retry logic collapses that ambiguity into 'it failed, try again', which is exactly the interpretation that causes duplicate side effects.

We've seen this with charges, with emails sent twice, with a ticket-status update that fired three times and tripped a downstream webhook each time. The blast radius is never the LLM. It's the tool call that mutated something real and got run again.

A timeout tells you the response didn't arrive, not that the work didn't happen.

Idempotent vs Non-Idempotent Is the Only Distinction That Matters First

Before you write a single line of retry logic, classify every tool the agent can call. Idempotent operations produce the same result no matter how many times you run them — a read, a GET, a `set_status(closed)` that's already closed. You can retry those freely. Non-idempotent operations mutate state cumulatively: create a record, send a message, move money, increment a counter. Retrying those without protection is how you get duplicates.

For non-idempotent calls you have two honest options. Make them idempotent with an idempotency key — a client-generated token the downstream service uses to dedupe, so a retry with the same key is a no-op that returns the original result. Stripe, most payment providers, and well-built internal APIs support this. If the downstream service can't dedupe, then the operation is genuinely unsafe to retry, and your retry policy for it should be: don't. Escalate or fail loud instead.

This classification isn't glamorous work and it's the step teams skip. It pairs directly with the refusal and permission boundaries you define before an agent can delete a record, the same audit that tells you which actions are irreversible tells you which are unsafe to retry.

  • →Idempotent (read, safe re-run): retry freely within budget.
  • →Non-idempotent + idempotency key supported: retry with the same key, downstream dedupes.
  • →Non-idempotent + no dedupe support: do not retry, escalate or fail loud.
  • →Unsure: treat as non-idempotent until proven otherwise.

Backoff Is Hiding Your Real Bug

Exponential backoff is the right tool for transient contention, rate limits, a momentarily overloaded service, a lock that'll clear in 200ms. It's the wrong tool for a systemic failure, and agents produce a lot of systemic failures dressed up as transient ones.

We audited an agent that was retrying a tool call three times with backoff on every request, adding seconds of latency and triple the token spend. The root cause wasn't a flaky service. The prompt was generating a malformed argument that the tool rejected with a 400, deterministically, every single time. Backoff turned an instant, loud, fixable prompt bug into a slow, expensive, silent one. The agent 'recovered' by eventually giving up — after paying three times to fail.

The rule: back off on retryable errors only. A 429 or a 503 is retryable. A 400, a 422, a schema-rejection, an auth failure, those will fail identically on retry, so retrying is pure waste. If your retry wrapper doesn't inspect the error class, it's not retry logic, it's a latency amplifier. This is also where budgeting latency across the whole agent path exposes the damage, silent retries are a top cause of a blown p99.

Four Responses to Four Failure Classes

The core mistake is treating 'something went wrong' as one event with one response. It's at least four, and they need different handling.

Retry when the failure is transient and the operation is safe to repeat, network blip, rate limit, idempotent call. Reprompt when the model produced bad output but the infrastructure is fine, malformed arguments, a refusal you can steer past. When you reprompt, you're changing the input, so give it its own bounded budget, because a model that got it wrong once often gets it wrong the same way twice. Escalate when the failure needs a human or a different system, an irreversible action the agent isn't allowed to take alone, or repeated failures that crossed a threshold. Fail loud when none of the above will help — an unretryable error on a non-idempotent call with no dedupe, or a budget exhausted.

The anti-pattern we see most is silent recovery that eats the signal. An agent that retries, reprompts, and then quietly returns a degraded answer teaches you nothing and hides the failure from your metrics, which is exactly the silent failure rate that never shows up in your ticket queue. Failing loud is a feature. It's how you find out what's actually broken.

  • →Retry: transient error + idempotent (or keyed) operation.
  • →Reprompt: valid infra, bad model output, with its own attempt budget, separate from retries.
  • →Escalate: irreversible action, or failure count over threshold, hand to a human or fallback path.
  • →Fail loud: unretryable + non-idempotent + no dedupe, or budget exhausted. Emit the error, don't swallow it.

The Guardrails We Wire In Before Go-Live

Three mechanisms turn this framework into something that survives contact with production. None of them are exotic; the failure is that they're treated as post-incident cleanup instead of day-one plumbing.

Idempotency keys on every non-idempotent tool call. Generate the key once per logical operation, not per attempt, and pass it through so that when a retry fires, the operation dedupes downstream instead of running twice. Generate a fresh key on retry and you've defeated the entire point.

Retry budgets, not retry loops. Cap total attempts per operation and total retries per request, and count reprompts separately. A per-request ceiling stops one bad task from spawning forty tool calls, the same tool-call sprawl that quietly triples your bill. When the budget is exhausted, you escalate or fail; you never retry into infinity.

A dead-letter queue for failures that survive the budget. Instead of dropping a permanently-failing operation or looping on it forever, park it with full context, the inputs, the error class, the attempt history, so a human or a batch process can resolve it. The DLQ is also your single best source of real edge cases for the eval set, which ties back to grading agents on their actions and side effects, not just their text.

Generate the idempotency key once per logical operation, not per attempt — a fresh key on retry defeats the entire point.

What's Still Hard

Two problems don't have clean answers yet. The first is the ambiguous timeout on a service that supports neither idempotency keys nor a status-check endpoint. You genuinely can't tell whether the work happened, and both retrying and not retrying can be wrong. The only real fix is upstream, pressure the service owner for a dedupe mechanism or a `get_operation_status` call, and until then you escalate rather than guess.

The second is deciding the retry budget itself. Set it too low and you fail on genuinely transient blips; too high and you boost latency and cost on systemic bugs. There's no universal number. We tune it from production traces — the distribution of how many attempts actually succeed on the second or third try tells you where the cliff is, and it's usually far lower than the default of three that everyone copies.

If you take one thing from this: classify your tools before you write your retry logic, and make failing loud the default rather than the exception. The agents that stay reliable in production aren't the ones that retry hardest. They're the ones that know exactly which failures they're allowed to retry, and refuse to guess on the rest.

Frequently asked questions

When should an AI agent retry a failed tool call versus escalate?

Retry only when the error is transient (429, 503, network blip) and the operation is idempotent or protected by an idempotency key. Escalate when the action is irreversible, unretryable, or has failed past a set threshold — retrying non-idempotent calls without dedupe just duplicates the side effect.

What's the difference between retrying and reprompting an LLM agent?

Retrying repeats the same tool call, expecting a different infrastructure result. Reprompting changes the model's input to fix bad output like a malformed argument or schema violation. They fail for different reasons and need separate, bounded budgets — a model that got it wrong once often repeats the same mistake.

How do idempotency keys prevent duplicate side effects in agents?

An idempotency key is a client-generated token, created once per logical operation, that the downstream service uses to dedupe. If a retry arrives with the same key, the service returns the original result instead of running the mutation again — so a retried charge or email doesn't fire twice. Generate the key per operation, never per attempt.

Why is exponential backoff sometimes the wrong choice for agents?

Backoff is for transient contention like rate limits. On a systemic failure — a prompt generating a malformed argument that the tool rejects deterministically — backoff just makes you wait longer and pay more tokens to fail identically. Inspect the error class first: only back off on genuinely retryable errors.

Free resource

Take the Operational Bottleneck Audit

Our Bottleneck Audit maps every tool call your agent makes, flags the non-idempotent ones running without protection, and finds where silent retries are burning latency and budget.

Ready to stop experimenting?

Find the retries quietly breaking your agent

We'll audit your agent's tool calls, retry paths, and side effects — and hand you the guardrails to wire in before the next duplicate-charge incident. Senior engineers only, 8–12 weeks.

Book a Discovery Call
More insights
Naive Retries Are Quietly Wrecking Your Production Agent | USQRD