Before your AI agent retries, give the action an identity

A practical design for tool calls that survive timeouts without duplicating the work.
An agent creates a support ticket. The service accepts it. The response disappears somewhere between the service and your worker.
The agent sees a timeout.
Its next move is reasonable: try again.
Now the customer has two tickets, two notifications, and two people investigating the same problem. The model understood the task. The integration still got it wrong.
In the previous post, we looked at what model-reported scores can and cannot tell you. Once a model can take actions, a different question becomes urgent: what happens when the system cannot tell whether an action already happened?
This is a design note for that boundary. The ticket service below is hypothetical; the architecture is a proposed pattern, not a claim about an existing Antler Labs implementation.
A timeout is an observation, not an outcome
A failed connection tells you what your caller observed. It does not necessarily tell you what the remote system did.
For a write operation, at least three possibilities matter: the request never arrived, it arrived but failed before making a change, or the change happened and the response was lost.
Treating all three as “failed, retry” collapses uncertainty into a decision that can create another side effect.
The Amazon Builders’ Library describes this ambiguity and the role of explicit caller-provided request identifiers. A repeated request can represent the same intent, while two requests with identical parameters can still represent two intentional operations.
That distinction is the starting point: identify the action you intend to perform, independently of the attempts to perform it.

Give the action a durable identity
For the hypothetical ticket workflow, separate three records:
| Record | What it identifies | Example |
|---|---|---|
| Run | One execution of the larger workflow | Investigate a customer report |
| Action | One intended change to an external system | Create the approved support ticket |
| Attempt | One effort to execute that action | First request, then a retry after a timeout |
A retry gets a new attempt ID and keeps the same action ID. A genuinely new ticket gets a new action ID.
Persist the action before sending the first request. That record can hold its tenant, operation, exact payload, approval reference, provider reference, status, and attempts. A process restart should reload this record rather than reconstruct the action from conversational memory.
The application should allocate and preserve the identity. Asking a model to remember which UUID it used turns a storage requirement into a language task.
There is a second boundary here: replaying the planning step must not silently allocate another action for the same workflow step. Resolve a stable step reference to its existing action record, or explicitly decide that the user requested a new operation. A fresh ID on every replay defeats the design before the HTTP request leaves your server.
Bind identity to intent
An action ID needs an immutable meaning.
For a ticket, that might include the destination workspace, requester, title, body, and any attachments. Your backend can normalize those fields, serialize them consistently, and store a digest alongside the original payload.
The digest helps detect changes. It should not become a global duplicate detector: two intentionally separate tickets can contain identical text.
The important rule is narrower:
Within the relevant tenant and operation scope, the same action ID must refer to the same intended change.
If a retry changes the destination or content, reject it as a mismatch. If the user wants a revision, create a new action and apply whatever approval policy that revision requires.
This gives the retry path a concrete contract instead of relying on whether two model outputs “look close enough.”
Carry that identity across the external boundary
A local action table is useful, but it cannot by itself prevent a remote service from doing the same work twice.
For providers that support idempotency keys, derive a stable provider key from the persisted action identity and reuse it for every attempt. Confirm the provider’s scope, retention window, parameter rules, and concurrent-request behavior.
Stripe’s API documentation is a concrete example: it describes replaying stored results for an idempotency key, rejecting mismatched parameters, and allowing keys to be pruned after at least 24 hours. Those details matter when designing recovery; provider deduplication is not an unlimited promise.
For our hypothetical ticket service, the recovery behavior should depend on its actual capabilities:
| Provider capability | Recovery approach |
|---|---|
| Documented idempotency support | Retry with the original key and payload, within the documented contract |
| Reliable lookup by a unique client reference | Reconcile against that reference; retry only if absence is authoritative and the original request cannot still complete |
| Neither capability | Keep the outcome unknown and route it to reconciliation or human review |
An eventually consistent search returning no ticket is not proof that creation failed.
Likewise, putting an external request inside a database transaction does not make the external service participate in that transaction. The worker can still crash after the provider commits and before your database records success.
Make “unknown” a real state
A useful action record distinguishes an unattempted operation from one whose outcome is uncertain.
For example, it can represent prepared, in_progress, succeeded, failed, and unknown. These are illustrative states, not a universal schema.
Use failed when you have evidence of a terminal failure. Use unknown when the action may have happened.
Recovery from unknown is a separate decision: consult the provider, use its supported idempotent retry path, or escalate. Do not quietly reset the action to prepared and let an ordinary worker send it again.
Concurrent workers need coordination too. Atomically claim work and use appropriate leases or locking. A lease expiring does not prove that the previous worker stopped or that its remote request failed. Any takeover still has to respect the provider boundary.
A transactional outbox can durably record that work needs dispatching. It still needs this same treatment when delivery is repeated.
Approval should cover the exact action

Suppose a user approves creating a ticket in the internal engineering workspace. After a timeout, the agent revises its plan and selects a customer-visible workspace instead.
The earlier approval should not authorize that change.
A practical approval record can bind the approver, action ID, payload digest, permission scope, and expiration. At execution, the backend checks that the persisted action still matches that approval and that current policy permits another attempt.
Reusing an approval for an unchanged retry can be reasonable if the policy allows it. A changed payload or expired authorization needs a fresh decision.
Idempotency and authorization answer different questions. One controls repetition of the same effect. The other controls whether the effect is permitted. A duplicate-free operation can still be unauthorized.
Test the moment after success
A happy-path demo will not exercise the boundary that matters most.
For this design, a small failure-focused test set would include:
- The provider commits the change, but the client never receives the response.
- The worker crashes after the remote commit and before saving the provider reference.
- Two workers try to execute the same action concurrently.
- The same action ID appears with a changed payload.
- Recovery starts after the provider’s deduplication window or the approval has expired.
With provider idempotency, verify that retries resolve to the same external effect under its documented guarantees. Without it, verify that an ambiguous outcome stays blocked for reconciliation instead of becoming an automatic repeat.
Also verify the legitimate opposite: two explicitly requested actions with identical payloads must remain possible.
The result should be an action history an operator can understand: what was intended, what was approved, which attempts ran, and what evidence establishes the outcome.
An agent can recover from uncertainty only if the surrounding system preserves it honestly. Before adding another retry, give the action an identity—and decide what that identity guarantees.
Part of the Antler Labs blog. Previous: What Verbalized Sampling buys you (and where scores lie).
Building an agent that needs to act reliably? Talk to Antler Labs.