SysDesignPrep.com
Study guide 155 of 183

Idempotency keys in practice

How to make retried API requests safe: what idempotency means, client-generated keys, storing keys with request fingerprints and responses, handling concurrent duplicates, expiry, keys across service boundaries, idempotent consumers, and the mistakes that still cause double charges.

Reading is half of it. See this used in a real interview: walk through Design a Payment System →

Networks fail in the worst possible place: after the server has done the work but before the client hears back. The client cannot tell whether the payment went through, so it retries, and without protection the customer is charged twice. Idempotency keys make retries safe: the same request, sent any number of times, has the effect of being sent once. They appear in every payment, ordering and booking design, and interviewers expect the implementation details, not just the term.

What idempotent means

An operation is idempotent if doing it twice has the same effect as doing it once.

  • GET, PUT (set to a value) and DELETE are naturally idempotent.
  • POST /payments (create a charge) and "add 10 to the balance" are not.

Idempotency keys turn non-idempotent operations into idempotent ones by recognising repeats.

The protocol

  1. The client generates a unique key (a UUID) for each logical operation, for example when the user presses "Pay", and sends it in a header: Idempotency-Key: 7c1e....
  2. The client reuses the same key for every retry of that operation, and a new key for a new operation.
  3. The server, on receiving a request:

- If the key is new: record it as in progress, execute, store the response, mark it complete. - If the key exists and is complete: return the stored response without executing again. - If the key exists and is in progress: return a conflict (409) or wait, so two concurrent duplicates do not both execute.

Payment providers such as Stripe implement exactly this.

Storage details

Store per key (scoped to the account or API client):

  • The key and the client or account id.
  • A fingerprint of the request (hash of method, path and body). If a retry with the same key has a different body, reject it: that is a client bug, not a retry.
  • Status: in progress, completed, failed.
  • The response status and body to replay.
  • Timestamps for expiry (keys are typically kept for 24 hours or more, longer than any client's retry window).

Atomicity is the hard part

The key record and the business effect must not diverge:

  • Same database: insert the key (with a unique constraint) and perform the business write in one transaction. Either both happen or neither. This is the most robust option.
  • Effect in another system (a call to a card network or another service): record the key and an intent first, then call the external system with an idempotency key of its own (derived from yours), then record the result. If the process crashes midway, a retry finds the intent and resumes or checks the outcome instead of starting fresh. See payments and ledgers.

Checking a key in Redis and then writing to Postgres leaves a window where a crash causes either a duplicate or a lost key. Use the database of record, or design for that window explicitly.

Concurrent duplicates

Mobile clients and impatient users can send the same request twice within milliseconds. The unique constraint on the key makes one insert win; the other sees the in-progress record and waits or returns 409. Never "check then insert" without a constraint.

Errors and retries

  • Do not store transient failures (timeouts, 503s from dependencies) as final responses; let the retry try again.
  • Do store deterministic outcomes (card declined, validation error), so retries get the same answer.
  • If an operation failed partway after side effects, the stored state must let the retry continue safely, which is where workflows help. See workflow orchestration.

Natural keys

Sometimes the business already provides a unique identity: an order id generated by the client, a (user, event, seat) tuple, a message id. A unique constraint on that natural key gives idempotency without a separate key table. Prefer it when available. See bookings and reservations.

Across service boundaries

Each hop needs its own protection:

  • Service A receives a request with key K, then calls service B with a key derived from K (for example K:charge), so A's retries do not create duplicates in B.
  • Events published to queues carry event ids; consumers deduplicate by them. See delivery semantics.
  • External providers (payment processors, email and SMS gateways) usually accept idempotency keys; use them.

Consumers of queues

Message consumers face the same problem: redelivery after a crash. Record processed message ids in the same transaction as the effect, or make the effect itself idempotent (upserts keyed by a deterministic id, conditional updates with versions). See Design a Notification System.

Common mistakes

  • Generating the key on the server: retries from the client then get new keys.
  • Generating a new key per retry in the client.
  • No request fingerprint, so a buggy client reuses keys for different operations.
  • Storing keys in a cache that can evict them before retries stop.
  • Expiring keys before the client's retry window ends.
  • Treating idempotency as a replacement for reconciliation: still reconcile money daily.

In the interview

"Clients send an Idempotency-Key per payment attempt; the server inserts the key with a request hash in the same transaction as the payment record, under a unique constraint; completed responses are replayed for retries; concurrent duplicates get a 409; calls to the card processor carry a derived key; keys expire after a few days." That covers Design a Payment System, and the same pattern applies to orders in Design DoorDash and purchases in Design Ticketmaster.

Checklist

  • Client-generated key per logical operation, reused on retries.
  • Key stored with request fingerprint, status and response.
  • Key and effect written atomically, with a unique constraint.
  • In-progress handling for concurrent duplicates.
  • Deterministic results stored; transient failures retried.
  • Derived keys for downstream calls; event ids for consumers.
  • Expiry longer than any retry window; reconciliation still in place.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.