Designing Idempotent Side-Effect Contracts for AI Agents
An AI agent calls send_refund. The payment provider accepts the request, but the response disappears 2026-9-28 09:0:20 Author: hackernoon.com(查看原文) 阅读量:2 收藏

An AI agent calls send_refund. The payment provider accepts the request, but the response disappears during a network timeout. The runtime sees an exception and retries. The second request also succeeds.

The trace now looks reassuring:

send_refund attempt=1 timeout
send_refund attempt=2 success
agent run success

The customer sees two refunds. This is a synthetic example, but the failure mode is ordinary distributed-systems behavior. The agent did not need a stranger prompt or a better model. It needed a tool contract that distinguished an unsuccessful response from an unsuccessful effect.

That distinction matters anywhere an agent can change the world: sending a message, creating a ticket, issuing a credit, booking a trip, deleting a record, or triggering a deployment. If the runtime treats every exception as permission to call again, “self-healing” becomes “self-duplicating.”

Developers often describe a tool with an input and output schema:

type RefundInput = {
  paymentId: string;
  amountCents: number;
  reason: string;
};

type RefundOutput = {
  refundId: string;
  status: "accepted";
};

That schema validates shape, not behavior. It does not tell the orchestrator:

  • whether the tool is read-only;
  • whether repeating it can create another effect;
  • whether it supports an idempotency key;
  • how to determine the outcome after a timeout;
  • which failures are safe to retry; or
  • whether human approval applies to one attempt or one business operation.

HTTP defines an idempotent method as one whose intended effect is the same after multiple identical requests, while also noting that a server may still log each request or create other non-idempotent side effects. That careful definition is useful for agent tools: the contract should be about the intended business effect, not merely about receiving one 200 response.

The Model Context Protocol now includes tool annotations such as readOnlyHint, destructiveHint, and idempotentHint. Those annotations help a client present and route tools, but the MCP maintainers explicitly describe them as untrusted hints, not security or correctness guarantees. A runtime still has to enforce its own policy.

Define a Side-Effect Contract

I prefer to make retry behavior part of the tool definition rather than scatter it across prompts and catch blocks.

type SideEffectClass =
  | "read_only"
  | "idempotent_write"
  | "deduplicated_write"
  | "non_repeatable_write";

type RetryPolicy = {
  maxAttempts: number;
  retryableCodes: readonly string[];
  requiresReconciliationAfterUnknown: boolean;
};

type ToolContract<I, O> = {
  name: string;
  sideEffect: SideEffectClass;
  validateInput(input: unknown): I;
  execute(input: I, context: ToolExecutionContext): Promise<O>;
  retry: RetryPolicy;
  reconcile?: (
    operationId: string,
    signal: AbortSignal,
  ) => Promise<ReconciliationResult<O>>;
};

The four classes force useful questions:

  • read_only: another call should not change business state.
  • idempotent_write: the target system guarantees one intended effect for equivalent calls.
  • deduplicated_write: safety depends on a stable operation or idempotency key.
  • non_repeatable_write: an automatic retry is forbidden unless reconciliation proves the first attempt did not commit.

The names are less important than making the policy executable.

Generate the Operation Identity Before the First Attempt

An idempotency key created inside each retry loop does not provide idempotency. Every attempt gets a new identity and therefore looks like a new operation.

Persist the operation before calling the external system:

type OperationState =
  | "prepared"
  | "in_flight"
  | "committed"
  | "rejected"
  | "outcome_unknown";

type OperationRecord = {
  operationId: string;
  runId: string;
  toolName: string;
  inputFingerprint: string;
  state: OperationState;
  providerReference?: string;
  attempts: number;
};

async function prepareRefund(
  runId: string,
  input: RefundInput,
): Promise<OperationRecord> {
  return operations.insertOnce({
    operationId: crypto.randomUUID(),
    runId,
    toolName: "send_refund",
    inputFingerprint: fingerprintValidatedInput(input),
    state: "prepared",
    attempts: 0,
  });
}

fingerprintValidatedInput should canonicalize a schema-approved, secret-free representation. A raw JSON.stringify is not a safe universal canonicalizer: key ordering, unsupported values, and sensitive input all need deliberate handling. RFC 8785 defines a JSON Canonicalization Scheme when interoperable hashing or signing is required.

Every attempt then reuses the same operationId:

async function executeRefund(
  operation: OperationRecord,
  input: RefundInput,
  signal: AbortSignal,
): Promise<RefundOutput> {
  await operations.transition(operation.operationId, "in_flight");

  try {
    const result = await payments.refund(input, {
      idempotencyKey: operation.operationId,
      signal,
    });

    await operations.markCommitted(operation.operationId, result.refundId);
    return result;
  } catch (error) {
    if (isDefinitiveRejection(error)) {
      await operations.transition(operation.operationId, "rejected");
      throw error;
    }

    await operations.transition(operation.operationId, "outcome_unknown");
    throw new UnknownOutcomeError(operation.operationId, { cause: error });
  }
}

The crucial line is not the idempotency header. It is the transition to outcome_unknown. A timeout tells you that observation failed. It does not tell you that the effect failed.

Unknown Outcomes Need Reconciliation, Not Hope

When the provider exposes an operation lookup, query it before allowing another write:

async function recoverRefund(
  operationId: string,
  signal: AbortSignal,
): Promise<"committed" | "not_found" | "still_unknown"> {
  const result = await payments.findRefundByIdempotencyKey(
    operationId,
    { signal },
  );

  if (result.status === "found") {
    await operations.markCommitted(operationId, result.refundId);
    return "committed";
  }

  if (result.status === "definitively_missing") {
    return "not_found";
  }

  return "still_unknown";
}

Only not_found can authorize a retry for a non-repeatable operation. still_unknown should pause, escalate, or schedule later reconciliation. It should not become a creative prompt asking the model what to do.

If the target system has no lookup and no deduplication support, that is part of the tool's contract. The safe behavior may be to require a human to inspect the external system. Automation cannot manufacture a guarantee that the dependency does not provide.

Approval Must Bind to the Operation

Suppose a user approves a $500 refund, the first attempt times out, and the runtime generates a second operation ID. Has the user approved the second potential refund?

No. Approval should bind to a canonical operation intent:

type ApprovalGrant = {
  approvalId: string;
  operationId: string;
  intentHash: string;
  approvedBy: string;
  expiresAt: string;
};

A retry with the same operation ID and unchanged intent can consume the same valid grant. A changed amount, recipient, destination, or operation identity requires a new approval. This prevents “retry” from becoming an accidental privilege expansion.

Use an Outbox When Local State and Dispatch Must Agree

Sometimes the agent records a decision in your database and publishes work to a queue. Writing the row and publishing the message as unrelated actions creates another ambiguity: the process can crash after either one.

The transactional outbox pattern places the business record and an outbox event in the same database transaction. A separate dispatcher publishes the event and records delivery attempts. Consumers still need deduplication because brokers and dispatchers commonly provide at-least-once delivery.

agent decision
    |
    v
[database transaction]
  operation row + outbox row
    |
    v
outbox dispatcher --may retry--> message broker
    |
    v
consumer deduplicates by operationId

This does not create magical exactly-once execution. It makes each boundary explicit and recoverable.

Trace the Effect Lifecycle

One tool.error=true field is not enough. Record the state machine without leaking raw payloads:

operation.prepared       id=op_73 tool=send_refund
operation.attempted      id=op_73 attempt=1
operation.outcome_unknown id=op_73 reason=deadline_exceeded
operation.reconciled     id=op_73 result=committed
agent.completed          id=run_12 outcome=refund_confirmed

Useful evidence includes the operation ID, sanitized intent fingerprint, state transitions, attempt count, policy decision, approval reference, reconciliation method, and provider reference. Never put secrets, full payment details, or uncontrolled user content into trace attributes.

For every side-effecting tool, I want written answers to these questions:

  1. What is the intended business effect?
  2. Is that effect naturally idempotent, deduplicated, or non-repeatable?
  3. Where is the stable operation identity created and stored?
  4. Which errors are definitive, retryable, or outcome-unknown?
  5. How does the runtime reconcile an unknown outcome?
  6. What changes invalidate approval?
  7. How are duplicate messages or callbacks handled?
  8. What evidence proves the final external state?

The model can propose an action. The runtime owns the effect.

A mature agent platform should never translate “I did not receive a response” directly into “do it again.” It should translate it into “the outcome is unknown—find out what happened.”

References


文章来源: https://hackernoon.com/designing-idempotent-side-effect-contracts-for-ai-agents?source=rss
如有侵权请联系:admin#unsafe.sh