Retries are one of the most common reliability mechanisms in automation. A connector times out, an API returns a transient error, a network request fails or a downstream platform is briefly unavailable, so the workflow tries again. That sounds simple until the first attempt actually completed but the caller never received the confirmation.
At that point the automation does not know whether it is recovering from a failed action or repeating a successful one. The technical failure and the business outcome are no longer the same thing.
A safe retry answers two questions: should this action run again, and how will we prove that it has not already happened?
The commercially meaningful problem
Duplicate work is not merely untidy data. It can create customer confusion, extra reconciliation, duplicate financial activity, inventory mismatches, repeated approvals and manual investigation. The larger the downstream consequence, the less acceptable a blind retry becomes.
This is why automation reliability has to be designed around the business outcome. Mellorca's earlier article on who watches the automations focuses on observability. Retry design is the next layer: once a failure is detected, recovery must not create a second problem.
Start by separating read operations from side effects
Some operations are naturally safer to repeat. Reading a customer record or checking a status generally does not change the world. Creating an invoice, sending a payment instruction, issuing a refund, creating an order, changing a stock quantity or dispatching a message does.
For every automated step, classify whether it has an external side effect. If it does, define what makes one business operation unique. That may be an order number, transaction reference, workflow instance, customer-and-period combination or another stable operation key.
Use idempotency where the platform supports it
Idempotency means repeating the same request does not create additional side effects. AWS documents this pattern directly: a request can include a client token so that a retry of the same operation returns the original result rather than performing the action again.
The principle is portable even when the platform uses different terminology. Before calling a side-effecting API, ask whether it supports an idempotency key, request identifier or equivalent duplicate-prevention mechanism. If it does, use a stable key derived from the business operation—not a new random value on every retry.
When the platform is not idempotent, make the workflow state-aware
Many business applications do not expose perfect idempotency. In that case, the workflow needs its own state check. Before creating something again, verify whether the intended outcome already exists.
For example, before creating an invoice, search for the business's unique invoice reference. Before sending a customer onboarding email, check whether the message event has already been recorded. Before creating a service ticket, check whether a ticket for the same triggering event already exists.
The check must use a reliable key. Matching on vague text, timestamps or names is usually too weak for business-critical duplicate prevention.
Do not treat every error as retryable
Retries are useful for temporary failures such as timeouts, rate limits and short service interruptions. They are usually harmful when the input is invalid, the user lacks permission, a required field is missing or a business rule rejects the action. Repeating a permanent error five more times only creates delay and noise.
Classify failures into at least three groups: retryable, non-retryable and uncertain. The uncertain group matters because it includes cases where the caller does not know whether the remote system completed the action. Those cases should trigger a state check before another write.
Use bounded retries and backoff
An automation should not hammer a failing service indefinitely. Define a maximum number of attempts, increasing delay between attempts where appropriate, and a point at which the workflow stops retrying and escalates.
Rate limits are a good example. Immediate repeated attempts can make the problem worse. A bounded retry policy gives the external platform time to recover and gives the business a predictable failure path.
Record enough evidence to reconstruct what happened
A retry policy is difficult to trust if nobody can tell which attempt succeeded. Record the workflow instance, operation key, attempt number, request time, response, resulting external identifier and final state. Do not log sensitive information unnecessarily, but preserve enough operational evidence to answer whether the business action happened.
This also supports the service-level thinking in business-critical automation service levels. Recovery time is only meaningful when recovery does not compromise data integrity.
A practical retry decision sequence
- Did the failed step have a side effect?
- Is the error genuinely transient?
- Does the target support idempotency or a unique request key?
- If not, can the workflow reliably check whether the outcome already exists?
- What is the maximum safe retry count and delay?
- What should happen when the retry limit is reached?
- What evidence will show which attempt produced the final outcome?
What better looks like
A mature automation does not equate resilience with persistence. It retries only when appropriate, identifies each business operation uniquely, prevents duplicate side effects, checks state when the result is uncertain and escalates cleanly when recovery cannot be completed safely.
The result is an automation that can fail without turning one incident into a reconciliation project.