# Do No Harm Twice: Idempotency, Explained Like You're New

In the replication post (#5), Section 4 left a grenade on the table: the primary dies mid-write, the client never got its "done," and now someone has to decide whether to retry the write. If the write actually landed, retrying charges the customer twice. If it didn't, not retrying loses the order. The network can't tell you which happened, because the *response* was lost, not necessarily the request, and guessing wrong costs money or customers. This post is the answer to that grenade: *idempotency*, the property that makes "just retry it" safe. It's the missing half of the resilience post's (#2) retry chapter, the quiet requirement inside the sharding post's (#3) sagas, and the reason payment systems sleep at night.

Here's what's covered: the double-charge and why retries are dangerous; what "idempotent" actually means, with the HTTP methods sorted correctly and the conditional-request machinery HTTP already gives you; idempotency keys, the full protocol in the right order (claim first, then act), the scope and fingerprint of a key, and the three error codes the emerging standard assigns; why exactly-once delivery is impossible, what Kafka's and SQS's "exactly-once" and "dedup" actually promise, and the outbox and inbox that close the gaps at both ends; the check-then-act race, where the atomic step lives, and why Redis is the wrong place to keep money's keys; why saga steps and compensations must be idempotent too; payments as the canonical case, with authorization and capture, refunds, chargebacks, and what reconciliation can and can't see on the day; the failure modes: key reuse, expiry, the key store's outage, partial failures inside one request, and the ordering trap; and the principal-level discipline: every mutation retryable, side effects you don't own, infrastructure as "make it so," dedup at real scale, testing with duplicate floods, and what all of it costs.

If you've never thought about what happens after a lost response, start at Section 1; the first two sections assume nothing, and every term is defined where it appears. Sections 3 through 8 are the machinery every backend engineer needs: keys, dedup, races, sagas, payments, failure modes. Section 9 is the judgment. The cheat sheet is at the end under *Idempotency, distilled*, and every diagram is described in the text around it, so nothing is lost on a screen reader.

---

## Section 1 — The double-charge

**In this section:** the incident that teaches the lesson, a retry that applied twice, and why "the network ate my response" is the most expensive sentence in distributed systems.

The story is a genre. A customer clicks "Pay $49." The request reaches the server, the charge succeeds, and then (a load balancer timeout, a deploy mid-request, a phone entering a tunnel) the response is lost on its way back. The app, helpfully, retries. The server sees a brand-new "Pay $49" request and charges again. The customer is charged $98 for a $49 purchase. Support refunds one charge, the customer leaves a one-star review with the word "scam" in it, and an engineer learns the lesson this post exists to teach: **in a distributed system, you cannot distinguish "the request failed" from "the response failed."** The request may have fully succeeded. You'll never know from the client's seat.

![Diagram: a client sends "Pay $49" — the server charges $49 (success) — the response is lost (red X on the return arrow, "response lost — client can't tell"). The client retries "Pay $49" — the server charges AGAIN — total $98. Below: "the client cannot distinguish a failed request from a failed response. Retrying blindly applies the effect twice."](https://res.cloudinary.com/i5oc7fbp/image/upload/v1790535802/idem/waaoevd2r3gnq5c0ip2w.png align="center")

This isn't a payments-only problem but every mutation over an unreliable network: the "submit order" tapped twice on a slow connection, the webhook (an HTTP call one service makes to another when something happens) delivered twice by a nervous sender, the saga step retried after a timeout (a saga being a multi-step transaction with undo steps; the sharding post's, #3, Section 7, and Section 6 here), the queue consumer that crashed *after* processing but *before* acknowledging. Anywhere a retry can happen, and retries are the resilience post's (#2) whole philosophy, applying the effect twice must be impossible, not merely unlikely.

Three words get used loosely here, and the rest of the post depends on keeping them apart. *Idempotency* is a property of an operation: doing it twice has the same effect as doing it once. *Deduplication* is a mechanism: recognizing that you've seen this request or message before and declining to process it again. *Exactly-once processing* is the outcome you want, and Section 4 will show that you get it by combining at-least-once delivery with one of the first two, never from the transport alone.

The naive fixes, and why they fail. "Disable the button after one click": the retry happens at the HTTP client, the queue, the load balancer; the button is the least of it. "Make the timeout longer": the failure mode doesn't care about your timeout values. "Check first, then act": two requests can both check, both see nothing, both act (Section 5 dissects this race properly). **The fix isn't preventing the second request. It's making the second request harmless.** That's idempotency.

One web-side pattern is worth its paragraph while buttons are on the table: *Post/Redirect/Get*. After a successful POST, answer with a redirect, so the browser's refresh re-issues the GET, not the POST. It fixes exactly one layer, the refresh button, but that's the layer your users actually touch.

When this clicks: the moment you internalize that the retry is not the bug. The retry is correct behavior in an unreliable world, and the non-idempotent handler is the bug. The rest of the post is implementation.

---

## Section 2 — What "idempotent" actually means

**Pin down the definition first; it's crisper than you think.** From the math to HTTP's method table, the conditional requests HTTP already provides, and the reframing trick that turns "do the thing" into "make it so."

The math first, because it's crisp: an operation *f* is idempotent if *f(f(x)) = f(x)*, applying it twice has the same effect as applying it once. "Set username to ana" is idempotent: set it twice, it's still ana. "Add $10 to the balance" is not: twice means +$20. Idempotency is a property of the *effect*, not the request. The same endpoint can be idempotent or not depending on what it does.

HTTP has known this since 1997, when RFC 2068 first defined the HTTP/1.1 methods (the current text is RFC 9110, from 2022). The *safe* methods, GET, HEAD, OPTIONS, and TRACE, are read-only and therefore idempotent by definition. PUT and DELETE are idempotent: replacing a resource twice or deleting it twice converges on the same state. POST is not: "create a new thing" applied twice creates two things. PATCH is not defined as idempotent either, because "apply this diff" can accumulate (a patch that says "append an item" appends twice), though a particular PATCH can be written to be. This is why REST style guides say updates should be PUT ("set the resource to this state") rather than POST ("do the thing"): **state-based operations ("make it so") are naturally idempotent; action-based operations ("do the increment") are naturally not.** When you have the choice, prefer "make it so."

![Diagram: two columns. Left, "naturally idempotent (make it so)": SET balance=100 → 100, applied twice → 100. PUT /users/9 {name: ana} twice → same. DELETE twice → gone, still gone. Right, "NOT idempotent (do the thing)": ADD 10 → 110, twice → 120. POST /charge $49 twice → $98. The rule: "state-based converges; action-based accumulates."](https://res.cloudinary.com/i5oc7fbp/image/upload/v1790535803/idem/v6m1zbcanucpfwwlz7jy.png align="center")

HTTP also ships the machinery for the most common "make it so" refinement, and it's underused. A *conditional request* carries an `If-Match` header with the `ETag` (a version tag) of the resource the client last saw; the server applies the PUT only if the resource still has that version and answers 412 Precondition Failed otherwise. That's optimistic concurrency control in two headers: a retry of the same PUT with the same `If-Match` either applies once (first attempt lost in transit) or fails cleanly (first attempt landed and bumped the version), and two users editing the same record can't silently overwrite each other. An API can even demand it, answering 428 Precondition Required to any update that arrives without a condition.

The reframing trick, and it's the sentence to remember from this section: **most non-idempotent operations can be rewritten as idempotent ones.** "Add $10" becomes "set balance to $110 *if it is currently $100*" (a conditional write). "Charge $49" becomes "create charge *with ID X*," and creating the same ID twice is a no-op (Section 3). "Process this event" becomes "process event *with sequence number N*, skipping N if already seen" (Section 4). The operation isn't idempotent or not; the *protocol around it* is.

Picture the team with a "deduct inventory" endpoint, action-based and non-idempotent, that keeps double-deducting on retries during deploys. The fix is not a lock or a queue but a change to the API, to "set inventory to N *with version V*": the client reads (N, V), computes the new N, and sends the write conditional on V still being current. Retries with the same V either apply once or fail cleanly on the version mismatch, never double-apply. Rather than adding idempotency machinery, they changed the verb from "do" to "make it so."

When to use: reach for the natural idempotency of PUT and state-based design first; it's free. Bring out the machinery (next section) for the operations that can't be reframed: payments, external side effects, anything where "create" is the business verb.

---

## Section 3 — The idempotency key

**The industry-standard machinery: the client-generated key.** The full protocol in the order that actually works, the three design decisions inside it, what a key is scoped to, and the emerging standard's error codes.

For operations that can't be reframed as "make it so," the standard answer is the *idempotency key*. Stripe made it famous (its engineers, Brandur Leach among them, wrote the design up in 2017), and the pattern is now in an IETF draft, "The Idempotency-Key HTTP Header Field," so the header name is settling on `Idempotency-Key`. The client generates a unique key per *logical* operation (a UUID, typically, up to 255 characters at Stripe) and sends it with the request. The server's protocol, in the order that survives concurrency:

1. **Claim.** Atomically record the key as *in progress* in the idempotency store. If the key is already there, don't execute anything: if the stored record has a response, return it; if it's still in progress, tell the caller to wait and retry (Section 5 explains why this step has to be atomic and has to come first).
2. **Execute** the operation.
3. **Store** the result against the key: the status code and body, whatever they were, and mark the record complete.
4. **Return** the response.

The retried request from Section 1 now plays out differently: the retry carries the same key, the server finds it in the store, and returns the stored "charged $49" response without touching the payment rail. **The second request is a cache hit, not a second charge.**

![Diagram: client sends POST /charge with Idempotency-Key: abc-123. Server checks the key store: "miss → execute charge → store key→response → return." The retry arrives with the same key: "HIT → return stored response, no re-execution." A note: "the store write and the effect must be atomic — or the effect idempotent itself (turtles all the way down, Section 5)."](https://res.cloudinary.com/i5oc7fbp/image/upload/v1790535805/idem/xxlkoo9qzbxumkan4ck6.png align="center")

Three design decisions, each load-bearing:

- **The client generates the key, per logical operation, not per HTTP attempt.** The retry must carry the *same* key, which means the key is created when the user *decides* ("I am placing this order"), not when the bytes go out. Where you can, derive the key from the business intent: `order:{cart_id}:checkout` beats a random UUID, because regenerating the intent (page reload, app restart) reproduces the same key for free, and the unique IDs post (#4) has the deterministic-ID version of the same idea. Server-generated keys work only in a two-step shape: the server hands the client an ID first (Stripe's PaymentIntent is created, then confirmed; a form can carry a one-time token), and the second step is idempotent on that ID. What can't work is the server inventing a key *inside* the request it's trying to protect, because the retry arrives before anyone knows it's a retry.
- **Store the response, not just a flag.** Returning the original response (with the original charge ID) makes the retry indistinguishable from the first attempt, and the client's world stays coherent. Stripe stores the status code and body of the first execution regardless of whether it succeeded or failed, so a retry of a request that got a 500 gets the same 500, which is the right answer: the client asked "what happened to *this* operation," not "please try again." Stripe also declines to store anything if validation failed before the endpoint started executing, so a rejected request can be corrected and re-sent under the same key.
- **Keys expire.** The store is not forever. Stripe prunes keys once they're at least 24 hours old, after which the same key is a new operation, and it accepts keys on POST only, since GET and DELETE are idempotent already. **The TTL is a contract**: it defines how long "retry safely" lasts. Pick it longer than your longest reasonable retry window (including a queue that got stuck for a day), and shorter than your storage budget's patience.

Two more things the protocol needs to define, and the draft standard names both. The *scope*: a key is unique within a namespace, usually the caller (the API key, the user, the tenant), so two customers who both happen to send `abc-123` don't collide, and one customer can't replay another's response. And the *fingerprint*: a hash of the request's meaningful parameters stored alongside the key, so that the same key arriving with a different body is rejected rather than answered with the first request's response. The draft's status codes are worth adopting as-is: 400 when a required key is missing, 422 when a key is reused with a different payload, and 409 when a request with the same key is still being processed.

Picture the mobile team that implements idempotency keys for its "place order" flow and puts the key generation in the HTTP layer, per *attempt*. Every retry gets a fresh key, so the server sees every retry as a new operation, and the double-orders continue, now with extra infrastructure. The fix is one line moved: generate the key when the order object is created on the device, and attach it to every attempt. **The key identifies the intent, not the attempt.** Get that wrong and the machinery is decoration.

When to use: any mutating API where the client might retry, which is any mutating API over a network. It's cheap (a table or key-value store with a TTL), standard (clients already know the header), and it composes with everything else in this post.

---

## Section 4 — Exactly-once is a lie (and at-least-once + dedup is the truth)

**One level deeper, into messaging, where "deliver exactly once" is provably impossible.** The architecture that works, what the vendors' "exactly-once" actually promises, and the outbox and inbox that close the gaps on both ends of the pipe.

Here's a result that surprises people the first time: **exactly-once message delivery is impossible in a distributed system.** The proof is short, and old: to guarantee the consumer processed a message exactly once, the broker must know the consumer *finished*, but the acknowledgment itself can be lost, so the broker can never be sure. It must either risk not delivering (at-most-once) or risk delivering twice (at-least-once). There is no third option. This is the *Two Generals problem*, stated in 1975 and named by Jim Gray in 1978: two parties on an unreliable channel can never reach certainty that both know a thing. Every queue, every webhook sender, every event bus you've ever used chose at-least-once, and either told you or didn't.

So the architecture the industry converged on: **at-least-once delivery plus idempotent consumers equals effectively-once processing.** The pipe may deliver twice; the consumer dedups. The dedup mechanism is a *deduplication store*: every processed message's ID is recorded (in a table, a cache, or, at scale, behind a Bloom filter, Section 9), and a message whose ID is already recorded is acknowledged without reprocessing.

![Diagram: a producer sends event E-42 to a queue. The queue delivers E-42 to the consumer — the consumer processes it, records "E-42 done," and its ack is LOST (red X). The queue redelivers E-42 (at-least-once). The consumer checks the dedup store: "E-42 already done → ack, skip processing." Below: "at-least-once delivery + idempotent consumer = effectively-once. The lie is in the pipe; the truth is at the edge."](https://res.cloudinary.com/i5oc7fbp/image/upload/v1790535806/idem/ik8adpnerfd7fu2klqro.png align="center")

The *dedup window* is the design parameter: how long do you remember processed IDs? Forever is correct and expensive; a TTL (24 hours, 7 days) is bounded and correct *as long as redeliveries can't arrive older than the window*. Size the window from the pipe's maximum redelivery age (the async post, #10, lists them: SQS keeps a message up to 14 days; a stuck consumer can hold one for as long as the visibility timeout allows) plus margin, not from a guess.

What the vendors actually promise is worth reading precisely, because the words "exactly-once" and "deduplication" appear in their docs with narrow meanings. Kafka's *idempotent producer* (the default since the 3.0 line, once a bug in its first releases was fixed, and only when no conflicting settings disable it) gives every producer a session ID and numbers each message per partition, so the broker discards a retried duplicate *from the same producer session*; a plain producer that restarts gets a new ID and the guarantee restarts with it (a transactional producer with a stable transactional ID carries it across restarts). Kafka *transactions* extend that to atomic writes across several partitions and to the read-process-write loop of a stream processor, which is what "exactly-once semantics" means there: exactly-once *inside Kafka*. The moment your consumer writes to a database or calls an API, you're back at the boundary and the dedup is yours. SQS's deduplication IDs exist only on FIFO queues, with a five-minute window; standard queues are at-least-once with no dedup at all, and the docs say so. Flink's checkpointed sinks are exactly-once for state Flink owns. Read every such claim as "exactly-once within the thing that's making the claim," and put your dedup at the edge anyway.

The version I've seen more than once: a team processing payment webhooks "handles" duplicates by doing nothing, for months, because duplicates are rare. Then the provider has an incident and replays six hours of webhooks, and the team credits hundreds of accounts twice before anyone notices. The postmortem's fix is a `processed_webhooks(event_id PRIMARY KEY)` table and a one-line check, the kind of fix that makes you angry it wasn't there from the start. Duplicates are rare until the day they're a flood. The dedup store is flood insurance. (Webhooks have a second requirement that isn't about duplicates: verifying the sender's signature before you trust the payload at all, which the security post, #12, covers.)

One more hole, on the *sending* side, because it pairs with everything above: the *dual-write problem*. Your service writes to its database *and* publishes to the queue: two writes, no shared transaction. Crash between them and the event is either lost (row committed, publish never happened) or phantom (publish happened, row rolled back). The canonical fix is the *transactional outbox*: in the same database transaction as your write, insert a row into an `outbox` table; a relay reads the outbox and publishes from it (or change data capture, from the replication post, #5, streams it straight from the database's log). The event can never be lost or phantom; at worst it's delivered twice, which is exactly what this section's consumer-side dedup is for. The receiving side's table has its matching name, the *inbox*: the consumer records the message ID in an `inbox` table *in the same transaction* as the work the message causes, so "processed" and "recorded as processed" can't come apart. Outbox on the way out, inbox on the way in, and the async post (#10) builds both.

When to use: every consumer of every at-least-once pipe, which is every pipe. If your consumer isn't idempotent, you don't have a consumer; you have a hope.

---

## Section 5 — The check-then-act race

**In this section:** the race that kills naive dedup. Two identical requests arriving at the same instant, both checking, both seeing nothing, both acting. Why the fix must be atomic, why the claim must come *before* the work, where the atomic step should live, and why Redis is the wrong home for money's keys.

Section 3's protocol has a hole if you run it in the wrong order, and it's the classic one: *check-then-act*. Request A checks the key store: miss. Request B checks the key store: miss. A executes and stores. B executes and stores. Two charges, one key. The window is microseconds wide and production *will* find it, because deploys, retries, and load balancers conspire to deliver duplicates simultaneously, not just one after another. A double-tap on a laggy phone sends two requests a few milliseconds apart, and a queue that redelivers on a timeout can hand the same message to two workers at once.

![Diagram: a timeline with two swim lanes, A and B. A: check → miss. B: check → miss (before A's store). A: execute + store. B: execute + store. Both charged. Red label: "the check-then-act race — the check and the act must be ONE atomic step." Below, the fix: a single "INSERT key ... IF NOT EXISTS" box — "the database's unique constraint decides the winner; the loser gets the stored response."](https://res.cloudinary.com/i5oc7fbp/image/upload/v1790535807/idem/l8nqwswyzdl9zapeofqs.png align="center")

The fix is to make check-and-claim one atomic step, and the database already has the primitive: the *unique constraint*. `INSERT INTO idempotency_keys (key, status) VALUES ($1, 'in_progress') ON CONFLICT DO NOTHING`, then look at whether your insert won. Or the equivalent conditional write in DynamoDB, or a Redis `SET key value NX` (set only if it doesn't exist). The database serializes the two inserts; one wins, one loses; the loser reads the winner's record and either returns its stored response or, if the winner is still in progress, answers 409 and lets the client retry in a moment. **The unique constraint is the arbiter.** The race is decided by the storage engine, not by your code.

Notice the ordering, because it's the part the diagrams get backward. The claim is inserted *before* the work runs, with a status of "in progress," and the response is filled in *after*. If you execute first and store afterward, two concurrent duplicates both execute in the gap. Claim, execute, complete. And the claim needs a plan for the crash in the middle: a record stuck "in progress" forever will block every retry of that key, so either give in-progress claims a short lease (after which a retry may take over and re-run, which requires the work itself to be safe to re-run) or have the request do its work inside the same database transaction as the claim, so a crash rolls both back together. That second design is the strongest one available. When the effect lives in the same database as the key (an order row, a ledger entry), put the key insert and the effect in *one transaction*: they commit together or not at all, and the response stored against the key is exactly the response that effect produced. When the effect lives elsewhere (a payment processor, an email), you get the recovery-point pattern from Section 8 instead.

Where the store lives matters as much as how it's used. Redis's `SET NX` is atomic, which is why it's so tempting, but by default Redis snapshots to disk every few minutes (the append-only log that would narrow that to about a second is off unless you turn it on) and replicates to its replicas asynchronously, so a crash or a failover can forget minutes of claimed keys, and every request in that window is now a fresh operation. For likes and notifications, that's a fine trade. For money, orders, and anything a customer will notice twice, keep the keys in the database that holds the effect, and let the database's durability guarantees (the replication post, #5, covers what those actually are) protect them.

This generalizes into a principle: **every dedup decision must bottom out in an atomic primitive**: a unique constraint, a conditional write, a compare-and-swap, a transaction. "Check in code, then act" is two steps pretending to be one, and under concurrency it's a race. When you review an idempotency implementation, the only question that matters is *where is the atomic step?* If nobody can point to it, the implementation is aspirational.

Picture the team that implements Section 3's protocol with a Redis `GET` then a `SET`: two round trips, a race window you could drive a truck through. It passes every test (tests don't do concurrency) and double-charges on the first real traffic spike. The fix is `SET key response NX EX 86400`, one command, atomic, TTL included, and then, because it's money, moving the whole thing into the orders database. The distance between "works" and "correct" was one command flag.

When to use: always. Any check-then-act in a dedup path gets the atomic treatment, no exceptions, no "it's unlikely."

---

## Section 6 — Sagas need it too

**The sharding post's sagas retry by design, which makes idempotency structural, not optional.** Every step, every compensation, every state transition.

The sharding post's (#3) Section 7 introduced the saga (Hector Garcia-Molina and Kenneth Salem, 1987): a distributed transaction as a sequence of local steps, each with a *compensating action* for rollback (book flight; on failure, cancel flight). Here's what that post didn't emphasize: sagas retry steps. If step 3 times out, the orchestrator retries step 3, because it can't know whether step 3 applied. And if the saga fails at step 4, the compensations run, possibly more than once, if a compensation times out. A saga is a retry machine that happens to be called a workflow, and that makes idempotency a structural requirement:

- **Each step must be idempotent.** Retried steps must not double-apply. The step is "reserve a seat *with reservation ID R*," and the reservation ID is the idempotency key from Section 3.
- **Each compensation must be idempotent.** "Cancel reservation R" twice must not cancel someone else's reservation or fail destructively. A compensation that errors on "already cancelled" wedges the rollback, so compensations treat "already undone" as success.
- **The saga's own state transitions must be idempotent.** "Mark step 3 complete" applied twice must not advance the saga twice, which means the orchestrator's state store needs the same atomic claim as everything else.

![Diagram: a saga — steps 1→2→3→4 left to right, each with a compensation arrow curving back below. Step 3 shows a timeout and retry: "retry with same reservation ID → no double-booking." The compensation for step 2 shows being run twice: "cancel R twice → second is a no-op." Label: "sagas retry by design — every step and every compensation is idempotent, or the workflow is a double-apply machine."](https://res.cloudinary.com/i5oc7fbp/image/upload/v1790535808/idem/skcb5llx6wo4xftrw9yi.png align="center")

Picture the travel-booking saga, flight then hotel then car, where the hotel step times out and retries, booking two rooms, and then the car step fails and the compensation cancels one room. The customer is charged for a room they never saw, and the "cancel booking" button in the app can't fix it because the app only knows about one. The root cause isn't the timeout or the failure; those are normal. It's non-idempotent steps in a retrying workflow. In a saga, "what if this runs twice" is not an edge case but the second line of the design doc.

When to use: any orchestrated multi-step workflow: sagas, workflow engines (Temporal and Cadence *activities* are the industrial version of a saga step, and Temporal's docs recommend that activities be idempotent for exactly that reason, because the engine will retry them), and webhook-chained integrations. If it has steps and retries, every step gets an idempotency story before it ships, and the async post (#10) covers the queues and outboxes that run the steps reliably.

---

## Section 7 — Payments: the canonical case

**The domain where idempotency isn't best practice but law.** Keys, the two-phase shape of a payment, ledger design, refunds and chargebacks, and reconciliation, with what it can and can't see on the day. The layered defense, because "twice" has a dollar sign.

Payments are the canonical case for a reason: a duplicate isn't a glitch, it's taking money twice for one purchase. So the payments industry built idempotency in layers, and the layering is worth studying because it's the template for any high-stakes mutation:

1. **Idempotency keys at the API** (Section 3). Stripe's `Idempotency-Key`, PayPal's `PayPal-Request-Id`, Adyen's `Idempotency-Key`, and even Authorize.Net's older "duplicate window" (which rejects an identical transaction within a two-minute window by default) are all versions of it. Same key, same charge object returned, never a second charge.
2. **The two-phase shape of a payment.** A card payment is usually *authorized* first (the bank holds the funds, nothing moves) and *captured* later (the money actually moves), which is why a retried "confirm" can be idempotent on the authorization's ID rather than on a fresh charge, and why a lost response on the capture step is recoverable: a second capture of the same authorization is refused rather than doubled (Stripe errors on an intent that's no longer capturable, and returns the original result if the retry carries the same idempotency key). Refunds get their own keys, because "refund $49" retried is the double-charge in reverse. And *chargebacks*, where the cardholder's bank reverses a payment, arrive as events you didn't initiate, which makes them a webhook-dedup problem (layer 5) and a ledger problem (layer 3), never a retry problem.
3. **Ledger design.** The money movement itself is recorded as immutable ledger entries, and the ledger carries a unique constraint on a *business* key, `(account, payment_intent_id)` say, rather than on the API's idempotency key, which expires in a day and belongs to one client. Even if every layer above fails, the ledger refuses the duplicate: Section 5's atomic arbiter, at the layer where money actually moves.
4. **Reconciliation.** An offline job continuously compares "what we think we charged" against "what the processor reports," catching anything the online path missed. Two clocks run here. Against the processor's *API* you can reconcile every few minutes (list today's charges, compare). Against *settlement*, the money actually arriving in your bank account, the report typically comes a day or more later (T+1 is the common shorthand), so the same job runs daily against that, and the two are not interchangeable. **Reconciliation is idempotency's backstop**, the admission that even good protocols deserve a second pair of eyes where money is involved.
5. **Webhooks with dedup** (Section 4). The payment *notifications* are at-least-once too: `event_id` primary keys on the receiving side, and signatures verified before anything else.

![Diagram: layered defense. Top: API with Idempotency-Key header → "retries return the stored charge." Middle: ledger with a unique constraint on the business key → "the money layer refuses duplicates atomically." Bottom: reconciliation job comparing internal ledger vs processor report → "catches anything the online path missed." Side: incoming webhooks → dedup store. Label: "five layers — no single layer is trusted alone with money."](https://res.cloudinary.com/i5oc7fbp/image/upload/v1790535809/idem/f9pbsqbay9y7zzms6vsq.png align="center")

Picture the idempotency review at a payments company, the one an engineer there might call the scariest meeting of the quarter: every new money-touching endpoint walks through where the key is, where the atomic step is, what the TTL is, what reconciliation sees, and what happens on the day the key store is down. Boring? Deeply. That's the point.

> **With money, idempotency isn't a feature. It's the code review.**

The numbers that discipline the design: key TTLs of at least 24 hours, and longer than your slowest pipe can redeliver, dedup windows sized to the processor's maximum replay age, API reconciliation every few minutes and settlement reconciliation daily. And the cultural rule: no money-touching endpoint ships without walking the layers.

When to use: verbatim, for payments. As a template (keys at the edge, an atomic constraint at the effect, reconciliation as backstop) for anything where "twice" has real-world cost: inventory, ticketing, access grants.

---

## Section 8 — Failure modes: keys, TTLs, and ordering

**In this section:** the ways idempotency machinery itself breaks. Reused keys, expired TTLs, key-store outages, the partial failure *inside* one request, and the ordering trap. See each one coming.

**Key reuse across operations.** The key identifies one logical operation (Section 3). Reuse a key for a *different* operation (a client bug, a key generator seeded wrong, a copy-pasted constant) and a naive server happily returns the first operation's response for the second. The customer is told "charged $49" for what should have been "refunded $49." Defense: the fingerprint from Section 3. Bind the key to the operation's parameters and reject mismatches (the IETF draft says 422; Stripe answers a 400 with an `idempotency_error`), which is exactly what Stripe does when the same key arrives with different parameters.

The most expensive version of the underlying sin wasn't a web API at all. In the last days of July 2012, Knight Capital deployed new trading code to seven of its eight servers, and on August 1 the market feature it supported went live. The deploy reused a flag that had once activated an old, discontinued order-routing feature, and on the eighth server, still running the old code, that flag meant the old thing. For about 45 minutes the firm's systems bought high and sold low across some 150 stocks, automatically, and the loss (about $440 million by Knight's own accounting, more than $460 million by the SEC's) ended the company as an independent firm. It's not an idempotency story strictly, but it is the same sin at a larger scale: one identifier that meant two things, and a system with no way to tell which one you meant.

**TTL expiry mid-retry.** Keys expire (24 hours, say). A retry storm or a stuck queue delivers the duplicate at hour 25, the key is gone, and the operation executes again. Defense: size the TTL from the *maximum* redelivery age of your slowest pipe, then add margin. Decide, too, what an expired key *means*: at Stripe a reused key after pruning is a new request; a stricter API can reject stale keys outright, which is safer for money and ruder to clients. And know that "forever" is a valid TTL when storage is cheap and the operation is dangerous; some ledger keys never expire.

**The key store is down.** Your idempotency store is now on the critical path of every mutation. If it's down, do you fail open (execute without dedup, risking duplicates) or fail closed (reject writes, an outage)? There is no comfortable answer, only the one you chose deliberately, per endpoint, in advance; the resilience post (#2) makes the same decision for every fallback. Most payment systems fail closed, because an outage beats double-charges. And if the keys live in the same database as the effect (Section 5), the question mostly disappears, because the store can't be down while the effect is up.

**Partial failure inside one request.** Real requests do more than one thing: charge the card, write the order, send the receipt. A crash after the charge and before the order row is the double-charge again, one level down, because a retry that starts from the top charges again. Brandur Leach's Stripe-style design handles this with *recovery points*: the key's record stores which phase completed, each phase that talks to the outside world runs in its own short transaction that records "phase 2 done, charge ID ch_123" before moving on, and a retry resumes at the recorded phase instead of restarting. The idempotency key stops being a boolean and becomes a small state machine, which is also how you make an operation resumable when the *client* gives up and comes back an hour later.

**The ordering trap.** Idempotency makes retries safe, but it doesn't order them: "set address to A" (key 1) and "set address to B" (key 2) can still apply in either order, because each is idempotent while the *sequence* isn't deterministic. If order matters, you need sequencing on top: per-client sequence numbers, a version on the resource (the `If-Match` from Section 2), or a single "make it so" carrying the full intended state. Idempotency guarantees "at most once per key." It says nothing about which key wins.

![Diagram: four panels, each a failure mode. 1) Key reused for a different operation → server returns the WRONG stored response — "bind keys to the operation fingerprint." 2) TTL expires at 24h, duplicate arrives at 25h → executes again — "size TTL from max redelivery age + margin." 3) Key store down → fork: "fail open (risk duplicates) or fail closed (outage) — choose deliberately." 4) Two idempotent writes race → either order possible — "idempotency ≠ ordering; add sequencing if order matters."](https://res.cloudinary.com/i5oc7fbp/image/upload/v1790535810/idem/rjtzdluidzfwzc6gamkg.png align="center")

Picture the team that fails *open* on a key-store outage, reasonable people choosing availability, and during a 20-minute outage a retry storm double-processes a batch of payouts. The money is recoverable (reconciliation, Section 7), but the week isn't. Their postmortem doesn't conclude "fail closed always." It concludes that the fail-open-or-closed choice is made per endpoint, in a design review, before the outage, and not by whoever's on call in the middle of the night.

When to revisit: every time you add a mutating endpoint, walk the five failure modes (the diagram shows four; the partial failure inside one request is the fifth). They're a checklist now.

---

## Section 9 — Going deep: the principal-level toolkit

**In this section:** the discipline. Every mutation retryable as a system property, the side effects you don't own, infrastructure as "make it so," dedup at true scale with honest numbers, idempotency as a platform capability, how to test it and watch it, and what it costs.

**The discipline, stated as a rule: every mutation in the system must be safe to retry.** Not "the important ones." All of them. The reasoning is the resilience post's: retries happen at layers you don't control (load balancers, service meshes, client SDKs, the user's thumb), so any non-idempotent mutation is a latent double-apply. Make it a code-review gate: a mutating endpoint without an idempotency story doesn't ship. The story can be "naturally idempotent (PUT semantics)" (Section 2), "idempotency key" (Section 3), or "conditional write" (Section 2's reframing), but "we'll add it later" is not a story. I ask for it in every design review now; the five minutes it costs is cheaper than the incident.

![Diagram: the code review gate — a mutating endpoint ships if and only if it has an idempotency story: naturally idempotent, idempotency key, or conditional write. "We'll add it later" is not a story and does not ship.](https://res.cloudinary.com/i5oc7fbp/image/upload/v1790535813/idem/q5acv57rpbyjjt17reda.png align="center")

**Side effects you don't own.** The hardest mutations to make idempotent are the ones that leave your systems: the email, the SMS, the push notification, the call to a partner API that has no idempotency key. You can't ask the email provider to un-send. Two techniques cover most of it. First, record *before* you act: insert "sending receipt for order 123" with a unique constraint on the order, in your own database, and only then call the provider; a retry finds the row and skips. That converts the external side effect into at-most-once, which is the right promise for a receipt (one missing email is a support ticket; three copies are a complaint) and the wrong one for a payment (where you'd rather have the recovery points from Section 8 and a reconciliation job). Second, when the third party offers *any* handle, use it: a message ID you supply, a "reference" field, a search by your own order number before you create. And when it offers nothing, put the call at the *end* of the request, after everything you can make atomic, so the window between "acted" and "recorded" is as small as you can make it.

**Infrastructure is the same discipline at a larger scale.** Declarative tooling is "make it so" at the level of whole systems: a Terraform plan or a Kubernetes manifest describes the desired state, and applying it twice converges instead of accumulating, which is exactly why those tools won over "run this script." Database migrations should be written the same way (`CREATE INDEX IF NOT EXISTS`, `ADD COLUMN IF NOT EXISTS`), because a migration that half-ran and gets re-run is Section 8's partial failure with a schema attached. And a deploy pipeline that can be re-triggered without redeploying twice is the same property again. If you've ever run `kubectl apply` a second time to be sure, you already believe in this section.

**Dedup at scale: Bloom filters, with real numbers.** Section 4's dedup store grows with every message, and at billions of events "have I seen this ID" becomes a storage problem of its own. The caching post's (#1) Section 9 introduced Bloom filters for exactly this shape: a probabilistic "definitely not seen / probably seen" check with no false negatives. The pattern: check the Bloom filter first; "definitely not seen" means process immediately (the common case, one memory lookup); "probably seen" means check the authoritative store. The sizes are worth knowing, because "billions of IDs in megabytes" is a myth: at a 1% false-positive rate a Bloom filter needs about 9.6 bits per item, so a hundred million IDs fit in roughly 120 MB and a billion in about 1.2 GB, still far smaller than the IDs themselves, and still in memory. The Bloom filter doesn't replace the dedup store; it keeps the hot path off it. False positives cost a store lookup and nothing else, because the store remains the arbiter, and Section 5's atomicity still lives there. One caveat: a plain Bloom filter can't forget, so a rolling window needs either a set of filters rotated by time or a counting variant.

![Diagram: Bloom filter dedup at scale — each arriving event first checks the Bloom filter; "definitely not seen" processes immediately with one memory lookup (the common case), while "probably seen" falls through to the authoritative dedup store, which decides new versus duplicate.](https://res.cloudinary.com/i5oc7fbp/image/upload/v1790535811/idem/snzczov5nri8nf3x06au.png align="center")

**Idempotency keys as a system property, not per-endpoint glue.** The mature shape: a shared library (or sidecar, or gateway plugin) that implements Section 3's protocol once, key extraction, atomic claim, fingerprint check, response replay, TTL, and the 400/409/422 responses, and every service gets it by configuration. The key store is shared infrastructure with its own SLO (service level objective, a reliability target it's held to). When idempotency is a platform capability, new endpoints are born idempotent; when it's per-endpoint glue, they're born whenever someone remembers. The one thing the platform can't do for you is the same-transaction design from Section 5, so the library should make "keys in your own database" the easy path, not the exotic one.

![Diagram: from per-endpoint glue — service A with hand-rolled keys, service B that forgot entirely, service C with a GET-then-SET race — to platform capability: a shared library, sidecar, or gateway plugin implementing key extraction, atomic check-and-store, response replay, and TTL against a shared key store with its own SLO.](https://res.cloudinary.com/i5oc7fbp/image/upload/v1790535812/idem/rwq3wusgft4mfy4f9ckh.png align="center")

**Test it, and watch it.** Idempotency bugs don't show up in unit tests, because unit tests don't do concurrency and don't lose responses. The tests that find them are the ones from the methodology post's (#14) "break it on purpose" section: send every mutating request twice, concurrently, and assert one effect; kill the worker after the effect and before the acknowledgment and assert the retry is harmless; replay yesterday's webhook stream into staging and count the side effects. On the observability side (#9), three numbers tell you whether the machinery is working: the replay hit rate (how often a request is answered from the store, which is your real duplicate rate, and it will surprise you), the 422 rate (a client misusing keys) and the 409 rate (concurrent duplicates, normal in small numbers and a misbehaving client in large ones), and the key store's latency and error rate, because it's on the critical path of every write now.

**What it costs.** The key store is a write on every mutation's critical path (latency, and a new dependency if it isn't your own database). Key storage is bounded by TTL but real. The protocol adds a header and a contract every client must learn. Response replay means storing response bodies. None of this is large, but it's load-bearing, which means it gets the testing, monitoring, and game days (rehearsed failure drills) of load-bearing things. The cost of idempotency is small and constant. The cost of its absence is rare and catastrophic. Price accordingly.

**The final reframe.** Idempotency is usually taught as a payments trick or an API nicety. It's bigger. It's the property that lets a system *act* in an unreliable world. Retries, failovers, redeliveries, saga replays: the entire distributed-systems toolkit assumes actions can be attempted again. Without idempotency, every one of those mechanisms is a loaded gun. With it, "just retry it," the three most relieving words in operations, is safe. **Design every mutation as if it will run twice. Because it will.**

---

## Idempotency, distilled

*For the skimmers and the revisitors: everything above, on one page.*

**Key numbers:**

| | |
|---|---|
| The fundamental ambiguity | A client can't distinguish a failed request from a failed response, ever (the Two Generals problem, 1975) |
| Exactly-once delivery | Impossible; at-least-once + dedup = effectively-once |
| HTTP's idempotent methods | GET, HEAD, OPTIONS, TRACE, PUT, DELETE (RFC 2068 in 1997, now RFC 9110); not POST, and not PATCH (RFC 5789) by definition |
| The protocol's order | Claim (atomic, "in progress") → execute → store the response → return |
| Standard error codes | 400 key missing, 409 same key still in progress, 422 same key with a different payload (IETF draft) |
| Stripe's contract | POST only; keys up to 255 chars; first status and body stored regardless of success; pruned after 24 h; parameter mismatch is an error |
| Atomic dedup primitive | Unique constraint / conditional write / `SET NX`: one step, decided by storage; same transaction as the effect when you can |
| Redis as a key store | Snapshots every few minutes by default (the once-a-second log is opt-in) and replicates asynchronously: fine for likes, not for money |
| Vendor "exactly-once" | Kafka: per producer session and inside Kafka; SQS dedup: FIFO queues only, 5-minute window; your edge still dedups |
| Dedup window sizing | The pipe's maximum redelivery age plus margin (SQS retains up to 14 days) |
| Bloom filter sizing | ~9.6 bits per item at 1% false positives: 100M IDs ≈ 120 MB, 1B ≈ 1.2 GB |
| Reconciliation clocks | Against the processor's API: minutes; against settlement: daily, typically T+1 |
| Knight Capital, 2012 | One repurposed flag on 1 of 8 servers; ~45 minutes; ~$440M (Knight) to $460M+ (SEC) |
| Fail-open vs fail-closed | Chosen per endpoint in design review, not by on-call during the outage |

**Every trade-off, in one table:**

| Decision | Chose | Over | Why |
|---|---|---|---|
| Retry safety | Make handlers idempotent | Prevent retries | Retries happen at layers you don't control; the handler is the fix |
| API shape | State-based (PUT, "make it so"), conditional on `If-Match` | Action-based (POST, "do the thing") | State-based converges on retry; action-based accumulates |
| Non-reframeable ops | Idempotency keys | Hope | Client-generated key per logical operation; server replays the stored response |
| Key identity | The intent (decision time) | The attempt (send time) | Retries carry the same key; server-generated keys only in a two-step shape |
| Key scope | Per caller, with a payload fingerprint | Global, key only | No cross-tenant collisions; a reused key with new parameters is a 422, not a replay |
| Stored on hit | The full status and body, success or failure | A boolean flag | The retry is indistinguishable from the original attempt |
| Protocol order | Claim, execute, complete | Execute, then store | Concurrent duplicates both execute in the gap |
| Where the keys live | The database that holds the effect, same transaction | A separate cache | Commit together or not at all; Redis can forget the last second |
| Check-then-act | Atomic (unique constraint / `SET NX`) | Two-step check then set | The race window is microseconds; production will find it |
| Messaging | At-least-once + idempotent consumer | An "exactly-once" pipe | Exactly-once delivery is impossible; dedup at the edge is the only design that works |
| Sending side | Transactional outbox (or CDC) | Write the row, then publish | Two writes, no transaction: a crash between them loses or fabricates an event |
| Receiving side | Inbox row in the same transaction as the work | Ack after processing and hope | "Processed" and "recorded as processed" can't come apart |
| Vendor claims | Scoped state guarantee + edge dedup | Trust "exactly-once" | Delivery to your code is still at-least-once |
| Dedup window | Max redelivery age + margin | A guess, or forever by default | Bounded storage that's correct; forever only where cheap and dangerous |
| Sagas | Idempotent steps, compensations, and state transitions | "Steps run once" | Sagas retry by design; non-idempotent steps are double-apply machines |
| Money | Layers: keys, auth-then-capture, ledger constraint on a business key, reconciliation, webhook dedup | One layer | No single layer is trusted alone with money |
| Multi-step requests | Recovery points per phase | Restart from the top | A crash between the charge and the order row is the double-charge again |
| Key-store outage | Fail open or closed, chosen per endpoint | Decide during the incident | Availability versus duplicates is a design-review decision |
| External side effects | Record before acting (at-most-once) or use any handle the third party offers | Act, then record | You can't un-send an email |
| Ordering | Sequencing or versions on top | Assume idempotency orders | At most once per key says nothing about which key wins |
| Scale | Bloom filter pre-check + authoritative store | A store lookup per message | Gigabytes in memory for billions of IDs; the store stays the arbiter |
| Adoption | Platform capability (library, sidecar, gateway) | Per-endpoint glue | New endpoints born idempotent versus born whenever someone remembers |
| Verification | Duplicate floods, kill-after-effect tests, replay hit-rate metrics | "It passed the unit tests" | Unit tests don't lose responses or run concurrently |

**Three ideas to take with you:**

1. **You cannot distinguish a failed request from a failed response — so stop trying.** The retry is correct behavior in an unreliable world. The bug is the non-idempotent handler, and the fix is making the second execution harmless, not preventing it.
2. **Every dedup decision bottoms out in an atomic step, taken before the work.** Unique constraint, conditional write, `SET NX`: somewhere, the storage engine decides the winner, and it decides before anything executes. If you can't point to that step in your implementation, it's aspirational.
3. **"Safe to retry" is a system property, not a feature.** Gate it in code review, build it as platform capability, test it with duplicate floods. Design every mutation as if it will run twice, because the network, the queue, the failover, and the saga all guarantee that it will.

---

## Further reading

- [Stripe, Idempotent requests](https://docs.stripe.com/api/idempotent_requests). The canonical idempotency-key API, including the 24-hour pruning and the store-the-result-even-on-failure rule; behind Sections 3 and 7.
- [Brandur Leach, Designing robust and predictable APIs with idempotency (Stripe, 2017)](https://stripe.com/blog/idempotency) and [Implementing Stripe-like idempotency keys in Postgres](https://brandur.org/idempotency-keys). The design rationale, and the atomic-phases and recovery-points pattern from Section 8.
- [IETF, The Idempotency-Key HTTP Header Field (draft)](https://datatracker.ietf.org/doc/html/draft-ietf-httpapi-idempotency-key-header). The standard-in-progress for the header, its scope, its fingerprint, and the 400/409/422 responses.
- [RFC 9110, HTTP Semantics, §9.2.2 Idempotent Methods](https://www.rfc-editor.org/rfc/rfc9110#section-9.2.2). The current definition of which methods are idempotent; behind Section 2.
- [Martin Kleppmann, *Designing Data-Intensive Applications*](https://dataintensive.net/). Exactly-once versus effectively-once, the Two Generals problem, transactions, and consensus; the deep end of Sections 4 and 9.
- [Garcia-Molina and Salem, Sagas (1987)](https://www.cs.cornell.edu/andru/cs711/2002fa/reading/sagas.pdf). The paper behind Section 6's compensating actions.
- [Apache Kafka documentation, Message Delivery Semantics](https://kafka.apache.org/43/design/design/#message-delivery-semantics). What the idempotent producer and transactions do and don't promise; Section 4's fine print, from the source.
- [Amazon SQS, Exactly-once processing in FIFO queues](https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/FIFO-queues-exactly-once-processing.html). The five-minute deduplication window, and why it's FIFO-only.
- [Chris Richardson, Transactional outbox pattern](https://microservices.io/patterns/data/transactional-outbox.html). The sending-side fix for the dual-write problem.
- [AWS Builders' Library, Timeouts, retries, and backoff with jitter](https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/timeouts-retries-and-backoff-with-jitter). The retry side of the contract; the resilience post's companion, and why Section 1's retries exist.
- [SEC, In the Matter of Knight Capital Americas LLC (2013)](https://www.sec.gov/files/litigation/admin/2013/34-70694.pdf). The order describing the repurposed flag and the 45 minutes; Section 8's cautionary tale, from the primary source.

---

## Where you'll meet this

This post is the answer to the replication post's (#5) Section 4 grenade, "did my write land?", and the missing half of the resilience post's (#2) retry chapter: retries are only safe when the retried thing is idempotent, which means those two posts and this one are one design decision. It was structural inside the sharding post's (#3) sagas (Section 7 there, Section 6 here): distributed transactions retry by design, so every step and compensation carries an idempotency story. The unique IDs post (#4) supplies the deterministic keys that make intent-derived idempotency free; the async post (#10) builds the outbox and inbox from Section 4; the security post (#12) verifies the webhooks before Section 4 dedups them; and the caching post's (#1) Bloom filters reappear here as the dedup-at-scale pre-check, the same probabilistic structure asked a different question. The [URL shortener](https://blueprintsofscale.hashnode.dev/from-laptop-to-billions-of-clicks-designing-a-url-shortener-in-11-steps) leaned on all of this in Step 6, where creating a short link had to survive a retried request without minting two keys. Next up in Core Concepts is what happens when the copy moves to the user's doorstep: CDNs and edge computing (#7).

---

## Keep exploring: the Core Concepts series

Every post in the series stands alone. Read them in any order.

- [From Laptop to Billions of Clicks: Designing a URL Shortener in 11 Steps](https://blueprintsofscale.hashnode.dev/from-laptop-to-billions-of-clicks-designing-a-url-shortener-in-11-steps) — the anchor: one design, every concept under load.
- [#1 The 100:1 Superpower: Caching](https://blueprintsofscale.hashnode.dev/the-100-1-superpower-caching-explained-like-you-re-new) — the fastest request is the one you never make.
- [#2 The Blast Radius: Surviving the Day Your Dependencies Fail](https://blueprintsofscale.hashnode.dev/the-blast-radius-surviving-the-day-your-dependencies-fail-explained-like-you-re-new) — staying up when everything you depend on goes down.
- [#3 Divide and Conquer: Sharding and Partitioning](https://blueprintsofscale.hashnode.dev/divide-and-conquer-sharding-and-partitioning-explained-like-you-re-new) — splitting one database into many without losing your mind.
- [#4 The Snowflake Problem: Unique IDs at Scale](https://blueprintsofscale.hashnode.dev/the-snowflake-problem-unique-ids-at-scale-explained-like-you-re-new) — naming things when millions are born every second.
- [#5 Copies of the Truth: Replication](https://blueprintsofscale.hashnode.dev/copies-of-the-truth-replication-explained-like-you-re-new) — keeping copies of your data that actually agree.
- **#6 Do No Harm Twice: Idempotency** — making "just retry it" safe. (this post)
- [#7 The Copy at the Doorstep: CDNs and Edge Computing](https://blueprintsofscale.hashnode.dev/the-copy-at-the-doorstep-cdns-and-edge-computing-explained-like-you-re-new) — serving from next door instead of across the ocean.
- [#8 The Bouncer's Math: Rate Limiting](https://blueprintsofscale.hashnode.dev/the-bouncer-s-math-rate-limiting-explained-like-you-re-new) — saying no politely, at scale.
- [#9 What Broke at 3 AM: Observability, p99, and Useful Alerts](https://blueprintsofscale.hashnode.dev/what-broke-at-3-am-observability-p99-and-useful-alerts-explained-like-you-re-new) — knowing what's wrong before your users tell you.
- [#10 Do It Later, On Purpose: Async Processing and Queues](https://blueprintsofscale.hashnode.dev/do-it-later-on-purpose-async-processing-and-queues-explained-like-you-re-new) — the work the user doesn't have to wait for.
- [#11 The Traffic Cop: Load Balancing](https://blueprintsofscale.hashnode.dev/the-traffic-cop-load-balancing-explained-like-you-re-new) — the box in every diagram nobody explains.
- [#12 Assume They're Already Knocking: Security and Abuse at Scale](https://blueprintsofscale.hashnode.dev/assume-they-re-already-knocking-security-and-abuse-at-scale-explained-like-you-re-new) — designing for the users who are designing against you.
- [#13 Do the Math First: Estimation for System Design](https://blueprintsofscale.hashnode.dev/do-the-math-first-estimation-for-system-design-explained-like-you-re-new) — the two minutes of arithmetic that choose the architecture.
- [#14 The Whiteboard Playbook: Taking On Any System Design Challenge](https://blueprintsofscale.hashnode.dev/the-whiteboard-playbook-taking-on-any-system-design-challenge-explained-like-you-re-new) — the whole series in forty-five minutes.

---

*Blueprints of Scale — Core Concepts #6. If this helped, the best thanks is a share with someone who's learning.*
