An AI agent calls send_refund. The payment provider accepts the request, but the response disappears during a network timeout. The runtime sees an exception and retries. The second request also succeeds.
The trace now looks reassuring:
send_refund attempt=1 timeout
send_refund attempt=2 success
agent run success
The customer sees two refunds. This is a synthetic example, but the failure mode is ordinary distributed-systems behavior. The agent did not need a stranger prompt or a better model. It needed a tool contract that distinguished an unsuccessful response from an unsuccessful effect.
That distinction matters anywhere an agent can change the world: sending a message, creating a ticket, issuing a credit, booking a trip, deleting a record, or triggering a deployment. If the runtime treats every exception as permission to call again, “self-healing” becomes “self-duplicating.”
Developers often describe a tool with an input and output schema:
type RefundInput = {
paymentId: string;
amountCents: number;
reason: string;
};
type RefundOutput = {
refundId: string;
status: "accepted";
};
That schema validates shape, not behavior. It does not tell the orchestrator:
HTTP defines an idempotent method as one whose intended effect is the same after multiple identical requests, while also noting that a server may still log each request or create other non-idempotent side effects. That careful definition is useful for agent tools: the contract should be about the intended business effect, not merely about receiving one 200 response.
The Model Context Protocol now includes tool annotations such as readOnlyHint, destructiveHint, and idempotentHint. Those annotations help a client present and route tools, but the MCP maintainers explicitly describe them as untrusted hints, not security or correctness guarantees. A runtime still has to enforce its own policy.
I prefer to make retry behavior part of the tool definition rather than scatter it across prompts and catch blocks.
type SideEffectClass =
| "read_only"
| "idempotent_write"
| "deduplicated_write"
| "non_repeatable_write";
type RetryPolicy = {
maxAttempts: number;
retryableCodes: readonly string[];
requiresReconciliationAfterUnknown: boolean;
};
type ToolContract<I, O> = {
name: string;
sideEffect: SideEffectClass;
validateInput(input: unknown): I;
execute(input: I, context: ToolExecutionContext): Promise<O>;
retry: RetryPolicy;
reconcile?: (
operationId: string,
signal: AbortSignal,
) => Promise<ReconciliationResult<O>>;
};
The four classes force useful questions:
read_only: another call should not change business state.idempotent_write: the target system guarantees one intended effect for equivalent calls.deduplicated_write: safety depends on a stable operation or idempotency key.non_repeatable_write: an automatic retry is forbidden unless reconciliation proves the first attempt did not commit.The names are less important than making the policy executable.
An idempotency key created inside each retry loop does not provide idempotency. Every attempt gets a new identity and therefore looks like a new operation.
Persist the operation before calling the external system:
type OperationState =
| "prepared"
| "in_flight"
| "committed"
| "rejected"
| "outcome_unknown";
type OperationRecord = {
operationId: string;
runId: string;
toolName: string;
inputFingerprint: string;
state: OperationState;
providerReference?: string;
attempts: number;
};
async function prepareRefund(
runId: string,
input: RefundInput,
): Promise<OperationRecord> {
return operations.insertOnce({
operationId: crypto.randomUUID(),
runId,
toolName: "send_refund",
inputFingerprint: fingerprintValidatedInput(input),
state: "prepared",
attempts: 0,
});
}
fingerprintValidatedInput should canonicalize a schema-approved, secret-free representation. A raw JSON.stringify is not a safe universal canonicalizer: key ordering, unsupported values, and sensitive input all need deliberate handling. RFC 8785 defines a JSON Canonicalization Scheme when interoperable hashing or signing is required.
Every attempt then reuses the same operationId:
async function executeRefund(
operation: OperationRecord,
input: RefundInput,
signal: AbortSignal,
): Promise<RefundOutput> {
await operations.transition(operation.operationId, "in_flight");
try {
const result = await payments.refund(input, {
idempotencyKey: operation.operationId,
signal,
});
await operations.markCommitted(operation.operationId, result.refundId);
return result;
} catch (error) {
if (isDefinitiveRejection(error)) {
await operations.transition(operation.operationId, "rejected");
throw error;
}
await operations.transition(operation.operationId, "outcome_unknown");
throw new UnknownOutcomeError(operation.operationId, { cause: error });
}
}
The crucial line is not the idempotency header. It is the transition to outcome_unknown. A timeout tells you that observation failed. It does not tell you that the effect failed.
When the provider exposes an operation lookup, query it before allowing another write:
async function recoverRefund(
operationId: string,
signal: AbortSignal,
): Promise<"committed" | "not_found" | "still_unknown"> {
const result = await payments.findRefundByIdempotencyKey(
operationId,
{ signal },
);
if (result.status === "found") {
await operations.markCommitted(operationId, result.refundId);
return "committed";
}
if (result.status === "definitively_missing") {
return "not_found";
}
return "still_unknown";
}
Only not_found can authorize a retry for a non-repeatable operation. still_unknown should pause, escalate, or schedule later reconciliation. It should not become a creative prompt asking the model what to do.
If the target system has no lookup and no deduplication support, that is part of the tool's contract. The safe behavior may be to require a human to inspect the external system. Automation cannot manufacture a guarantee that the dependency does not provide.
Suppose a user approves a $500 refund, the first attempt times out, and the runtime generates a second operation ID. Has the user approved the second potential refund?
No. Approval should bind to a canonical operation intent:
type ApprovalGrant = {
approvalId: string;
operationId: string;
intentHash: string;
approvedBy: string;
expiresAt: string;
};
A retry with the same operation ID and unchanged intent can consume the same valid grant. A changed amount, recipient, destination, or operation identity requires a new approval. This prevents “retry” from becoming an accidental privilege expansion.
Sometimes the agent records a decision in your database and publishes work to a queue. Writing the row and publishing the message as unrelated actions creates another ambiguity: the process can crash after either one.
The transactional outbox pattern places the business record and an outbox event in the same database transaction. A separate dispatcher publishes the event and records delivery attempts. Consumers still need deduplication because brokers and dispatchers commonly provide at-least-once delivery.
agent decision
|
v
[database transaction]
operation row + outbox row
|
v
outbox dispatcher --may retry--> message broker
|
v
consumer deduplicates by operationId
This does not create magical exactly-once execution. It makes each boundary explicit and recoverable.
One tool.error=true field is not enough. Record the state machine without leaking raw payloads:
operation.prepared id=op_73 tool=send_refund
operation.attempted id=op_73 attempt=1
operation.outcome_unknown id=op_73 reason=deadline_exceeded
operation.reconciled id=op_73 result=committed
agent.completed id=run_12 outcome=refund_confirmed
Useful evidence includes the operation ID, sanitized intent fingerprint, state transitions, attempt count, policy decision, approval reference, reconciliation method, and provider reference. Never put secrets, full payment details, or uncontrolled user content into trace attributes.
For every side-effecting tool, I want written answers to these questions:
The model can propose an action. The runtime owns the effect.
A mature agent platform should never translate “I did not receive a response” directly into “do it again.” It should translate it into “the outcome is unknown—find out what happened.”