Retries Are Not a Safety Plan
Retries make an agent look calm right up until it repeats the same bad tool call five times. I have watched a workflow burn through time and money while doing exactly the wrong thing with impressive consistency.
A retry is useful when the failure is probably temporary: a connection reset, a rate limit, or a provider that was briefly unavailable. It is not a safety strategy for a malformed argument, a missing permission, or a tool that keeps rejecting the same request.
The agent needs a small budget and a reason for each attempt. I want to know whether it is retrying the transport, changing the input, or blindly replaying the same action because the loop has no better idea.
Tool calls also need to be safe to repeat. If a call creates a record, sends a message, or charges money, idempotency matters more than how persuasive the model sounds when it asks to try again.
When the budget runs out, the system should show the failure clearly and leave enough evidence for somebody to decide what to do. Quietly hiding the problem behind another attempt only makes the eventual recovery harder.
I still use retries. I just treat them as one small part of reliability, alongside timeouts, typed errors, idempotency, tracing, and a clear hand-off when the system has reached the edge of what it can safely decide.