Edition 2: Prevent duplicate AI actions after a timeout

A timeout does not tell you whether an external action failed. The email may have been accepted, the CRM record may exist, or the payment request may have reached its provider even though your workflow received no response. If an agent or automation treats that silence as permission to try again, one logical task can become two customer-facing actions.

This lesson is for a founder or small operations team connecting an AI workflow to email, CRM, ticketing, invoicing, or payments. You will build a local action gateway that gives each approved business action a stable identity, records uncertainty durably, checks the provider before retrying, and routes unprovable outcomes to a person. The completed artifact is a Python 3.12 program with a restart test, a deliberately lost response, an unavailable-readback case, exact expected output, and a small cost calculation.

The model is not the safety boundary here. The gateway is. A model may propose an action, but deterministic code decides whether that action can cross the external boundary.

Why this problem earned an issue

Two dated problem signals support the reader need.

First, a LangGraph issue opened July 28, 2026 reports that long-running or retried executions can re-invoke tools and duplicate side effects after restarts or timeouts. The discussion identifies a particularly important crash window: the external write succeeds, then the worker stops before saving the receipt. This is a public issue report and design discussion, not a measurement of how often every framework fails.

Second, a LangChain forum question posted February 17, 2026 documents confusing repeat execution around an interrupt and resume flow. The reported duplicate AI messages are not the same as duplicate payments or emails, but the question shows that workflow replay behavior is difficult for implementers to reason about.

A preprint submitted July 31, 2026 separately studies non-atomic tool failures in a controlled simulation and tests postcondition checks, verification before retry, and idempotency keys. It supports the method, but it is a simulated study and does not establish field prevalence, customer savings, or a universal implementation.

I also sampled guidance returned for the question “How do I prevent an AI agent from repeating an external action after a timeout?” on October 11, 2026. AWS guidance on retry-safe APIs explains client request identifiers and reconciliation. Stripe's idempotency documentation explains how repeated requests with the same key reuse the first result. An Orkes guide covers idempotency in distributed workflows. This lesson adds a small-team lab that makes the unknown middle state visible, kills the response after the write, restarts from disk, refuses a blind retry when readback is unavailable, and includes review effort in the result. That is the bounded gap addressed here. It is not a claim that no other implementation exists.

Outcome, boundary, and acceptance rule

The business example is a reviewed welcome-message workflow. Three approved lead records reach the action gateway. The provider is simulated locally, so no message is sent and no account is touched.

Before running anything, use these acceptance criteria:

  1. Three logical actions create exactly three provider records.
  2. A lost response leaves the local action UNKNOWN, not failed.
  3. After a process restart, provider readback settles the unknown action without a second provider write.
  4. If readback is unavailable, the action becomes REVIEW and the gateway does not retry it.
  5. Reusing one action key with different payload data is rejected.
  6. The program prints CHECK PASS only when two actions are settled, one is held for review, and no duplicate provider action exists.

The naive baseline intentionally “sends” the same logical welcome twice after a simulated timeout. Its duplicate count is one. The safe workflow must reduce that synthetic count to zero. This is a correctness test on three invented records, not a production reliability estimate.

Prerequisites, permissions, time, and cost

  • Python 3.12. The example was executed with Python 3.12.14.
  • A terminal and permission to create files in one disposable directory.
  • No package installation, API key, customer data, cloud account, or network permission.
  • About 30 minutes to read, run, inspect the databases, and try the exercise.
  • Local exercise API cost: $0.00. Disk and staff time are not literally free. The sample assigns four minutes of review at an illustrative loaded rate of $30 per hour, producing a $2.00 review cost.

For a real integration, the narrowest useful permission is usually create plus readback for one resource type. Do not give a workflow account-wide administration merely because it needs to create a ticket or draft a message. Keep payment capture, irreversible deletion, and customer promises behind separate approval until their recovery procedures have been tested.

The four states

PREPARED means the stable identity and payload hash are saved before any external call. UNKNOWN means the provider may have committed, but the response was lost. SETTLED means a provider response or later readback produced a receipt. REVIEW means code cannot prove whether a second attempt is safe.

That distinction matters because “error” collapses two different facts: “the provider proved it did nothing” and “the caller does not know what happened.” Only the first fact can authorize a retry. The second requires reconciliation or a person.

The stable action key expresses business intent, not an execution attempt. welcome:lead-202:v1 means one version of one approved welcome action. Attempt number two must reuse that key. A new key would incorrectly describe a new intention.

The payload hash prevents another mistake: reusing the same key for altered content or a different recipient. Provider-side idempotency is valuable only when the key and payload retain their meaning.

Step 1: Create a clean lab

Run:

mkdir retry-safe-lab
cd retry-safe-lab
python3 --version

Save the following complete program as retry_safe_actions.py.

#!/usr/bin/env python3
"""A local lab for retry-safe business actions. Python 3.12, standard library only."""
from __future__ import annotations

import hashlib
import json
import sqlite3
import sys
from dataclasses import dataclass
from decimal import Decimal
from pathlib import Path


class ResponseLost(RuntimeError):
    """The provider committed the action, but the caller did not receive the response."""


class ReadbackUnavailable(RuntimeError):
    """The provider cannot currently confirm whether an action exists."""


def canonical(value: dict) -> str:
    return json.dumps(value, sort_keys=True, separators=(",", ":"))


def payload_hash(value: dict) -> str:
    return hashlib.sha256(canonical(value).encode()).hexdigest()


@dataclass(frozen=True)
class Action:
    key: str
    recipient_ref: str
    template: str
    simulate: str = "ok"

    @property
    def payload(self) -> dict:
        return {"recipient_ref": self.recipient_ref, "template": self.template}


class DemoProvider:
    """A fake provider with durable idempotency and readback by business key."""

    def __init__(self, db_path: Path):
        self.db = sqlite3.connect(db_path)
        self.db.row_factory = sqlite3.Row
        self.db.execute(
            """CREATE TABLE IF NOT EXISTS effects (
               action_key TEXT PRIMARY KEY,
               payload_hash TEXT NOT NULL,
               receipt TEXT NOT NULL
            )"""
        )

    def perform(self, action: Action) -> str:
        digest = payload_hash(action.payload)
        found = self.db.execute(
            "SELECT payload_hash, receipt FROM effects WHERE action_key = ?", (action.key,)
        ).fetchone()
        if found:
            if found["payload_hash"] != digest:
                raise ValueError("same action key was reused with different data")
            return found["receipt"]
        receipt = "msg_" + hashlib.sha256(action.key.encode()).hexdigest()[:10]
        self.db.execute(
            "INSERT INTO effects(action_key, payload_hash, receipt) VALUES (?, ?, ?)",
            (action.key, digest, receipt),
        )
        self.db.commit()
        if action.simulate in {"lost", "unreadable"}:
            raise ResponseLost("response disappeared after provider commit")
        return receipt

    def readback(self, action: Action) -> str | None:
        if action.simulate == "unreadable":
            raise ReadbackUnavailable("provider readback is unavailable")
        row = self.db.execute(
            "SELECT payload_hash, receipt FROM effects WHERE action_key = ?", (action.key,)
        ).fetchone()
        if not row:
            return None
        if row["payload_hash"] != payload_hash(action.payload):
            raise ValueError("provider data does not match the prepared action")
        return row["receipt"]

    def count(self) -> int:
        return self.db.execute("SELECT COUNT(*) FROM effects").fetchone()[0]


class ActionLedger:
    def __init__(self, db_path: Path):
        self.db = sqlite3.connect(db_path)
        self.db.row_factory = sqlite3.Row
        self.db.execute(
            """CREATE TABLE IF NOT EXISTS actions (
               action_key TEXT PRIMARY KEY,
               payload_hash TEXT NOT NULL,
               state TEXT NOT NULL,
               attempts INTEGER NOT NULL DEFAULT 0,
               receipt TEXT,
               note TEXT NOT NULL DEFAULT ''
            )"""
        )

    def prepare(self, action: Action) -> sqlite3.Row:
        digest = payload_hash(action.payload)
        self.db.execute(
            "INSERT OR IGNORE INTO actions(action_key, payload_hash, state) VALUES (?, ?, 'PREPARED')",
            (action.key, digest),
        )
        self.db.commit()
        row = self.get(action.key)
        if row["payload_hash"] != digest:
            raise ValueError("same action key was reused with different data")
        return row

    def get(self, key: str) -> sqlite3.Row:
        row = self.db.execute("SELECT * FROM actions WHERE action_key = ?", (key,)).fetchone()
        if row is None:
            raise KeyError(key)
        return row

    def update(self, key: str, state: str, *, receipt: str | None = None, note: str = "") -> None:
        self.db.execute(
            "UPDATE actions SET state = ?, receipt = COALESCE(?, receipt), note = ? WHERE action_key = ?",
            (state, receipt, note, key),
        )
        self.db.commit()

    def add_attempt(self, key: str) -> None:
        self.db.execute("UPDATE actions SET attempts = attempts + 1 WHERE action_key = ?", (key,))
        self.db.commit()

    def counts(self) -> dict[str, int]:
        rows = self.db.execute("SELECT state, COUNT(*) count FROM actions GROUP BY state").fetchall()
        return {row["state"]: row["count"] for row in rows}


def execute(action: Action, ledger: ActionLedger, provider: DemoProvider) -> dict:
    row = ledger.prepare(action)
    if row["state"] == "SETTLED":
        return {"key": action.key, "state": "SETTLED", "attempts": row["attempts"], "path": "cached"}

    if row["state"] == "UNKNOWN":
        try:
            receipt = provider.readback(action)
        except ReadbackUnavailable:
            ledger.update(action.key, "REVIEW", note="readback unavailable; no retry allowed")
            row = ledger.get(action.key)
            return {"key": action.key, "state": row["state"], "attempts": row["attempts"], "path": "manual"}
        if receipt:
            ledger.update(action.key, "SETTLED", receipt=receipt, note="confirmed by readback")
            row = ledger.get(action.key)
            return {"key": action.key, "state": row["state"], "attempts": row["attempts"], "path": "reconciled"}
        ledger.update(action.key, "PREPARED", note="readback proved no action")

    row = ledger.get(action.key)
    if row["state"] == "REVIEW":
        return {"key": action.key, "state": "REVIEW", "attempts": row["attempts"], "path": "manual"}

    ledger.add_attempt(action.key)
    try:
        receipt = provider.perform(action)
    except ResponseLost:
        ledger.update(action.key, "UNKNOWN", note="provider may have committed")
        row = ledger.get(action.key)
        return {"key": action.key, "state": row["state"], "attempts": row["attempts"], "path": "hold"}
    ledger.update(action.key, "SETTLED", receipt=receipt, note="provider response received")
    row = ledger.get(action.key)
    return {"key": action.key, "state": row["state"], "attempts": row["attempts"], "path": "sent"}


def naive_baseline() -> dict:
    external_effects = []
    for attempt in range(2):
        external_effects.append({"attempt": attempt + 1, "recipient_ref": "lead-202"})
    return {"attempts": 2, "external_actions": len(external_effects), "duplicates": 1}


def main(folder: Path) -> None:
    actions = [
        Action("welcome:lead-201:v1", "lead-201", "welcome-v1", "ok"),
        Action("welcome:lead-202:v1", "lead-202", "welcome-v1", "lost"),
        Action("welcome:lead-203:v1", "lead-203", "welcome-v1", "unreadable"),
    ]
    ledger = ActionLedger(folder / "operator.sqlite3")
    provider = DemoProvider(folder / "provider.sqlite3")

    print("BASELINE " + canonical(naive_baseline()))
    first = [execute(action, ledger, provider) for action in actions]
    print("FIRST_RUN " + canonical(first))

    # A new ledger object represents a worker restart using the same durable database.
    restarted = ActionLedger(folder / "operator.sqlite3")
    recovered = [execute(action, restarted, provider) for action in actions]
    print("RESTART " + canonical(recovered))

    counts = restarted.counts()
    review_minutes = counts.get("REVIEW", 0) * 4
    review_cost = Decimal(review_minutes) / Decimal(60) * Decimal("30.00")
    summary = {
        "duplicates": 0,
        "exercise_api_cost_usd": "0.00",
        "logical_actions": len(actions),
        "provider_actions": provider.count(),
        "review_cost_usd": f"{review_cost:.2f}",
        "review_minutes": review_minutes,
        "states": counts,
    }
    print("SUMMARY " + canonical(summary))
    passed = provider.count() == 3 and counts == {"REVIEW": 1, "SETTLED": 2}
    print("CHECK " + ("PASS" if passed else "FAIL"))
    if not passed:
        raise SystemExit(1)


if __name__ == "__main__":
    target = Path(sys.argv[1]) if len(sys.argv) > 1 else Path(".")
    target.mkdir(parents=True, exist_ok=True)
    main(target)

The program uses two SQLite files. operator.sqlite3 is your action ledger. provider.sqlite3 represents the external system. Keeping them separate preserves the important failure boundary: the provider can commit even when the operator fails to save the returned receipt.

Step 2: Run the hostile timeout and restart test

From the clean directory, run:

python3 retry_safe_actions.py run

The optional run argument names a subdirectory that will hold both databases. Do not reuse that directory when comparing your first output with the expected block.

Expected complete stdout:

BASELINE {"attempts":2,"duplicates":1,"external_actions":2}
FIRST_RUN [{"attempts":1,"key":"welcome:lead-201:v1","path":"sent","state":"SETTLED"},{"attempts":1,"key":"welcome:lead-202:v1","path":"hold","state":"UNKNOWN"},{"attempts":1,"key":"welcome:lead-203:v1","path":"hold","state":"UNKNOWN"}]
RESTART [{"attempts":1,"key":"welcome:lead-201:v1","path":"cached","state":"SETTLED"},{"attempts":1,"key":"welcome:lead-202:v1","path":"reconciled","state":"SETTLED"},{"attempts":1,"key":"welcome:lead-203:v1","path":"manual","state":"REVIEW"}]
SUMMARY {"duplicates":0,"exercise_api_cost_usd":"0.00","logical_actions":3,"provider_actions":3,"review_cost_usd":"2.00","review_minutes":4,"states":{"REVIEW":1,"SETTLED":2}}
CHECK PASS

The first lead settles normally. The second provider write succeeds but its response disappears. After the restart, readback finds the existing receipt and settles it with one attempt. The third write also succeeds, but provider readback is unavailable, so the gateway stops in REVIEW. It does not assume failure and it does not manufacture a second key.

The measured findings are limited to this execution: three synthetic logical actions, three simulated provider records, zero duplicates in the safe path, one item requiring review, and exact stdout matching the block above. The baseline duplicate is deliberately programmed. No live email, CRM, payment, customer, deployment, or vendor bill was involved.

Step 3: Connect a real provider at one documented boundary

Replace only DemoProvider.perform and DemoProvider.readback. Keep the ledger and state transitions outside model-controlled text.

For perform, send the stable action key using the provider's documented idempotency field. Stripe, for example, accepts an Idempotency-Key on POST requests and documents retention and parameter-mismatch behavior. Other providers use different names and guarantees. Verify the exact API version you call.

For readback, retrieve the result by that same key, a provider metadata field, or a unique business identifier. Return the provider receipt only after checking that recipient, action type, amount or template version matches the prepared payload hash. A list search that can miss recently written data is not proof that nothing happened.

Never put an email address, customer name, card token, or message body in the idempotency key. Use an opaque business reference. Store sensitive payloads only where your retention and access policy permits them. The lab stores sanitized references and a hash, not message content.

The export boundary is the local actions table. A production adapter can expose these nonsecret fields to an incident queue or CSV: action key, state, attempts, receipt reference, note, creation time, and last update. Do not export raw credentials or customer bodies merely to make a dashboard convenient.

If the provider has no idempotency or readback contract, do not wrap a dangerous create call in automatic retries. Use a draft-only action, a unique constraint you control, or a manual queue. For money movement, use the provider's official preview and idempotency mechanisms and reconcile against authoritative transaction records.

Step 4: Measure the human and system cost

The script treats one unresolved item as four review minutes. At an illustrative loaded rate of $30 per hour:

4 minutes / 60 × $30.00 = $2.00

Replace both inputs with your measurements. Record review time from ordinary runs and incidents separately. Add provider fees, database and queue costs, engineering maintenance, and the cost of delayed work. Do not count a prevented duplicate as saved revenue unless you have an auditable counterfactual.

Useful operating measures are:

  • logical actions requested;
  • external calls attempted;
  • provider effects found;
  • unknown and review states;
  • readback success and age;
  • reviewer minutes;
  • duplicate effects confirmed by authoritative records;
  • completed actions that meet their business acceptance rule.

The ratio between attempts and completed logical actions is often more informative than raw API call count. Pair it with review effort, because a workflow that avoids duplicates by sending every case to a person has not completed the job efficiently.

Failure cases and troubleshooting

Symptom: the same key raises “different data.” The workflow changed the recipient, template, amount, or other protected field after preparation. Create a reviewed new logical action and key. Do not weaken the hash check.

Symptom: an action stays UNKNOWN. The provider has not exposed an authoritative receipt yet. Wait only within a documented visibility window, then route to review. Never translate a timeout into “not sent.”

Symptom: readback returns several candidates. Your key is not unique at the provider boundary. Stop automation, compare authoritative records, and strengthen provider metadata or your controlled database constraint.

Symptom: local database is locked. Keep the transaction that claims or updates an action short. Do not hold it open during a network request. For concurrent production workers, move the ledger to a database with well-understood row locking and a unique constraint on the action key.

Symptom: the provider accepts an idempotency key but a duplicate appears. Check whether the retry occurred after the provider's retention window, whether a new key was generated, and whether the operation is actually covered by the contract. Preserve request IDs and provider receipts for support.

Rollback and manual fallback. Disable automatic dispatch, preserve the ledger, and export UNKNOWN and REVIEW rows to an authorized operator. The operator checks the provider's authoritative record, records the receipt if found, or explicitly approves a new action if nonexecution can be proved. Do not delete uncertain rows to make the queue look clean.

Completion check and independent exercise

You are done when a clean run matches the expected stdout exactly, CHECK PASS appears, the two SQLite files exist, and you can explain why the third action was not retried.

For an independent exercise, add a fourth action and simulate “provider rejected before execution.” Add a PROVED_NOT_EXECUTED outcome from readback, move the ledger back to PREPARED, and permit exactly one new attempt with the same action key. Write the acceptance rule before editing: four logical actions, four provider effects, no duplicate effects, and no remaining unknown state. Then add the negative test that the same key with a changed template raises ValueError.

The next learning step is to put the adapter behind an approval queue and add timestamps, actor identity, permission scope, and a redacted incident export. That turns retry safety into an auditable operating workflow rather than a hidden helper function.

Evidence labels and limits

  • Documented: Stripe and AWS describe idempotency keys, retry behavior, and reconciliation considerations in their own systems.
  • Source-inspected report: The LangGraph issue and forum thread document specific implementer concerns. They do not establish prevalence.
  • Research-reported: The July 31 preprint reports a controlled simulation. This lesson does not claim independent reproduction of the paper.
  • Synthetic: Lead references, provider behavior, timeout, unavailable readback, review time, and hourly rate in this lab.
  • Measured here: Python 3.12.14 execution, three logical actions, three provider records, two settled states, one review state, zero safe-path duplicates, and exact stdout match.
  • Not established: Production uptime, customer savings, inbox delivery, payment safety for a specific provider, search demand volume, traffic, or revenue.

Continue with the seven-day reviewed lead-intake pilot to place this gateway inside a bounded workflow, and use the human-review time method to measure whether recovery work changes the business case.

Published by Primus Vekuh. This lesson has no paid placement or affiliate links.