Make Your Automation Idempotent Before You Make It Smart
Most broken AI automations aren't broken because the model is dumb - they're broken because every retry does the work twice. Here's the contract I use for idempotent pipelines, with schemas, hashing rules and an audit checklist.
Pavel Duglas
AI Automation & MVP Architect
The single most expensive bug I’ve shipped in an automation had nothing to do with AI. A client’s invoicing workflow timed out on the last step, the platform retried it, and 340 invoices went out twice. The model was fine. The prompt was fine. The problem was that my pipeline had no concept of “I already did this.”
Since then I treat idempotency as a design decision I make on day one, before I write a single prompt. Especially now, when half the pipeline is an LLM that returns something slightly different every time you call it. If you’re building agents, scrapers, Telegram bots or n8n flows that touch real systems, this is the least glamorous and highest-leverage thing you can fix this week.
Idempotency is a contract, not a column
A lot of developers hear “idempotency” and think “I’ll add an idempotency_key column with a unique index.” That’s a mechanism, not a contract. The contract has three parts, and if you don’t decide all three explicitly, you don’t have idempotency - you have a race condition with a unique index on it.
1. Identity. What exactly counts as “the same work”? Same webhook delivery? Same order? Same order in the same state? This is a business question, not a technical one, and getting it wrong is how you end up either duplicating invoices or silently swallowing legitimate second orders from the same customer.
2. Window. How long do you remember? Forever? 24 hours? Until the next successful sync? A dedupe cache with a 5-minute TTL is fine for webhook storms and useless for a nightly job that reruns after a two-day outage.
3. Response. What happens on a duplicate? Do you return the original result, return an error, or quietly no-op? An agent calling your tool needs to know the difference between “created” and “already existed, here’s the original” - otherwise it will invent a retry loop.
Write those three answers down in a comment above the handler. I’m serious. Half the arguments I’ve had on client projects disappeared once identity was written in one sentence.
Step 1: Deterministic work IDs, not random UUIDs
The caller generating a fresh UUID per attempt is useless. The ID must be derivable from the work itself, so that two attempts at the same work produce the same ID.
Two strategies, and you’ll usually want both:
Business key - when the source system gives you a stable identifier:
work_id = sha256("invoice:v1:" + shopify_order_id)
Content hash - when it doesn’t (scraped pages, form submissions, inbound emails):
import { createHash } from "node:crypto";
function workId(kind, version, payload) {
const canonical = JSON.stringify(payload, Object.keys(payload).sort());
return createHash("sha256")
.update(`${kind}:${version}:${canonical}`)
.digest("hex");
}
Three rules I learned the hard way:
- Canonicalize before hashing. Sort keys, trim strings, normalize numbers and timestamps. Otherwise
{a:1,b:2}and{b:2,a:1}are “different work.” - Strip volatile fields.
received_at,request_id,trace_idwill destroy every hash. Hash only the fields that define the work. - Version the prefix. When your logic changes and you want everything reprocessed, bump
v1→v2. That single decision has saved me from writing migration scripts more times than I can count.
Step 2: Keep an effect ledger, not a dedupe set
A seen_ids set answers “did I start this?” It doesn’t answer “did I finish it, and what came out?” You need the second answer, because the dangerous window is exactly when a job crashed mid-flight.
create table work_log (
work_id text primary key,
kind text not null,
state text not null check (state in ('running','done','failed')),
attempt int not null default 1,
result jsonb,
external_ref text, -- id returned by the third-party system
locked_until timestamptz,
created_at timestamptz not null default now(),
updated_at timestamptz not null default now()
);
The claim is a single atomic statement - no read-then-write, no application-level locking:
insert into work_log (work_id, kind, state, locked_until)
values ($1, $2, 'running', now() + interval '10 minutes')
on conflict (work_id) do update
set state = 'running',
attempt = work_log.attempt + 1,
locked_until = now() + interval '10 minutes',
updated_at = now()
where work_log.state = 'failed'
or work_log.locked_until < now()
returning *;
If you get a row back, you own the work. If you get zero rows, someone else is doing it or it’s already done - read the stored result and return that. This handles the three cases that actually happen in production: concurrent duplicate deliveries, retries after a crash, and retries after a real failure.
The external_ref column is the underrated one. It’s your proof that the effect landed in Stripe, Sheets, or the CRM, and it’s what lets a reconciliation job clean up without guessing.
Step 3: With LLMs, separate the decision from the effect
Here’s where AI pipelines get their own flavor of this problem. Your model is non-deterministic. Retry a step and you may get a different classification, a different extracted amount, a different tool call. If the retry happens after a side effect, you now have two different effects from “the same” work.
So I split every LLM step into two phases.
Phase 1 - decide, and cache the decision. Key the cache on a hash of everything that influences output: model name, prompt template version, temperature, tool schema, and the input payload.
decision_key = sha256(model + prompt_v + str(temperature) + tools_hash + input_hash)
Store the parsed output. On retry you replay the same decision instead of rolling the dice again. Bonus: this is also the cheapest LLM cost optimization there is, and it makes your pipeline debuggable - you can see exactly what the model said on attempt 1.
Phase 2 - apply, idempotently. The apply step takes the cached decision plus the work_id and performs the effect through the ledger above. Model output proposes; only the apply step commits.
This also gives you a free kill switch: a plan you can inspect before it executes. When a client asks “can we approve the agent’s actions before they go out,” the answer is already yes, because the decision is already a stored artifact.
Step 4: The outside world is your problem too
Your ledger protects your database. It does not protect the third-party API you’re calling. Three tiers, in order of preference:
Tier 1 - the API supports idempotency keys. Stripe, and a growing number of others. Send your work_id as the key. Done. Never generate a random one per attempt; that defeats the entire purpose.
Tier 2 - the API has a natural unique field you control. Set an external reference, a slug, an SKU, a custom field. Then “create” becomes “upsert by that field.” Most CRMs and billing tools support this if you look.
Tier 3 - nothing. Telegram sendMessage, most webhooks, most legacy endpoints. Here you do check-then-act and accept it’s imperfect: search for an existing record by a fingerprint you embedded (a short hash in a note field, a hidden line in the message), and only create if absent. Then narrow the race window by making the call last in the transaction sequence and recording external_ref immediately after.
For outbound messaging specifically: keep a sent_messages table keyed by (chat_id, content_hash) with a TTL. It won’t be perfect under a network partition, but it stops the classic failure where a queue redelivery spams 400 users at 3am.
Step 5: Retries, windows, and the dead letter you actually read
Assume at-least-once delivery everywhere. Your queue, your webhook provider, your n8n error branch, your cron that overlaps with itself - all of them will double-fire eventually.
- Bound the attempts. Exponential backoff with jitter, max 5 tries, then dead letter. Infinite retries on a poison payload will burn your API quota by morning.
- Make the dead letter visible. A Telegram channel with
work_id,kind, last error and a one-click replay command. If nobody sees it, you don’t have a dead letter queue, you have a data loss feature. - Overlap protection on schedules. A cron that runs every 5 minutes and sometimes takes 7 needs a lock, not hope.
locked_untilin the ledger already gives you one. - Reconciliation beats prevention. Once a day, compare your ledger to the target system and report drift. This is the only thing that catches Tier-3 races, and it’s usually 40 lines of code.
The scraping and BAS angle
Long-running browser automation has the same disease in a different costume. A BAS script that dies at item 8,400 of 10,000 and restarts from zero isn’t just slow - it re-submits forms, re-sends messages, re-burns proxy traffic and account limits.
The fix is the same shape: give every item a deterministic ID (usually the source URL or listing ID, normalized - strip tracking params and sort query strings), persist processed IDs outside the run, and checkpoint a cursor per run. On restart, claim only unprocessed items. And keep run_id separate from work_id - a new run is not new work.
One extra rule for scrapers: normalize before you hash. I’ve seen dedupe fail completely because ?utm_source= variants made every URL unique. Same page, five “different” work items, five duplicate rows.
A 30-minute audit you can run today
Open your most business-critical automation and answer these out loud:
- If this runs twice with the same input, what breaks? Name the specific side effect.
- Where does the deterministic ID come from? If the answer is “the platform’s execution ID,” you have no idempotency.
- If it crashes between the LLM call and the write, what happens on retry? Do you re-prompt?
- Which external calls have real idempotency keys, which have natural unique fields, and which are hope-based?
- After 5 failures, where does the payload go, and who sees it?
- Can you replay a single failed item without rerunning the whole batch?
Every “I don’t know” is a production incident with a date on it you just haven’t read yet.
What I deliberately skip
I don’t make read-only steps idempotent - fetching, summarizing for a dashboard, enrichment that writes to a cache. Re-running them costs tokens, not trust. I also skip full ledgers on internal tools with one user who can see the duplicate and delete it in two seconds.
The rule I use: if a duplicate touches money, a customer’s inbox, or an external system’s state, it gets the full contract. Everything else gets a retry and a shrug. Being smart about where you don’t need this is what keeps the pattern from turning into ceremony.
FAQ
Isn't a unique index enough for idempotency?
A unique index stops duplicate rows, but it doesn't tell you what to return on a duplicate, and it doesn't help when the effect lives outside your database. If your job crashes after calling Stripe but before inserting the row, the index never sees the conflict - the money moved anyway. That's why you need a ledger that records state and the external reference, not just a constraint.
How do I make an LLM step idempotent when the model output changes every run?
Split it in two. First, the decision: call the model once and cache the parsed result under a key that hashes the model name, prompt version, temperature, tool schema and input. Second, the effect: apply the cached decision through your idempotent write path. On retry you replay the stored decision instead of re-prompting, so the effect stays consistent - and you get cheaper token usage and a debuggable audit trail for free.
What do I do when the third-party API has no idempotency key support?
Fall back in tiers. If there's a field you control that's unique (external reference, slug, SKU, custom field), turn create into upsert on that field. If there's nothing at all, embed a short fingerprint hash in a note or message body, search for it before writing, record the returned external ID immediately, and add a daily reconciliation job that compares your ledger to the target system and reports drift. It's not perfect, but it catches the duplicates that check-then-act misses.
Related articles
Done for you
I will build an AI agent for a real task
With tools, memory and logs, so it works in production and not only in a demo.
from $1,500 · 1 to 2 weeks