Designing webhooks
How to deliver webhooks reliably: event payloads, signing and verification, retries with backoff, ordering, idempotency on the receiver, per-customer isolation, and replay, as Stripe and GitHub do it.
Reading is half of it. See this used in a real interview: walk through Design a Payment System →
A webhook is your system calling a customer’s server when something happens: a payment succeeded, a repository was pushed, an order shipped. It sounds like one HTTP request. In practice the customer’s endpoint is slow, down, misconfigured or malicious, and you have to deliver millions of events reliably anyway. Designing webhooks well is a frequent interview topic for payments and platform companies.
Event model
Publish events, not internal state changes: payment_intent.succeeded, invoice.paid. Each event has:
{
"id": "evt_1Q2w3E", // unique, for deduplication
"type": "payment_intent.succeeded",
"created": 1727689000,
"data": { "object": { …the resource at the time of the event… } },
"api_version": "2026-09-30"
}Two choices to make deliberately:
- Fat payloads (the full resource) let receivers act without calling you back, but data in the event may be stale by the time it arrives.
- Thin payloads (just the id and type) force receivers to fetch the current state from your API, which is always fresh and avoids leaking data to a misconfigured endpoint.
Many platforms send fat payloads and tell receivers to fetch the resource when correctness matters.
Delivery pipeline
- The event is written in the same transaction as the change that caused it (a transactional outbox), so an event exists if and only if the change committed.
- A dispatcher fans the event out to every endpoint subscribed to that event type, creating one delivery per endpoint.
- Delivery workers POST to the endpoint with a short timeout (5 to 10 s) and record the result.
- Failures are retried; after the last attempt the delivery is marked failed and the endpoint may be disabled.
Deliveries are background jobs; see background jobs and task queues.
Retries
- Retry on timeouts, connection errors and 5xx responses; treat 2xx as success; treat most 4xx as permanent failures (except 429, which means slow down).
- Exponential backoff with jitter over a long window: Stripe retries for up to three days. A customer’s overnight outage should not lose events.
- Disable endpoints that fail for days, and notify the customer by email.
- Never let retries for one bad endpoint delay deliveries to everyone else.
Isolation between customers
One customer’s slow endpoint (30-second responses) can tie up a shared worker pool and delay every other customer’s webhooks. Prevent it with:
- Per-endpoint queues or concurrency limits, so each endpoint gets at most a few in-flight requests.
- Short timeouts, and moving slow endpoints to a separate, lower-priority pool.
- Rate limits per endpoint, matched to what the receiver can handle.
Security
- Sign every request: compute an HMAC of the timestamp and body with a per-endpoint secret, and send it in a header (
Signature: t=1727689000,v1=…). Receivers verify it and reject requests older than a few minutes to stop replay attacks. - HTTPS only.
- Protect yourself from SSRF: customers choose the URL, so block private and internal address ranges (169.254.x.x, 10.x, localhost) and re-check after DNS resolution, or a customer can make your workers call your own internal services.
- Support secret rotation with two valid secrets during the switch.
Ordering and duplicates
Webhook delivery is at least once and not ordered: a retry of event 1 can arrive after event 2. Tell receivers to:
- Deduplicate by event id (store processed ids).
- Not rely on order: compare timestamps or versions on the resource, or fetch the current state from the API when an event arrives.
Guaranteeing order per resource is possible (deliver events for one resource sequentially), but one failing event then blocks all later ones for that resource, so most platforms do not.
Observability for customers
Customers need to debug their side, so provide:
- A log of deliveries per endpoint with request, response code, latency and attempts.
- Manual replay of an event or a time range.
- A test event button and a CLI that forwards events to a local machine.
- An events API to list recent events, so receivers can reconcile after an outage instead of relying on webhooks alone.
Receiving webhooks (the other side)
If your design consumes webhooks (from a payment provider, say): verify the signature, respond 2xx quickly, put the event on a queue and process it asynchronously, deduplicate by event id, and periodically reconcile against the provider’s API, because webhooks can be missed. See Design a Payment System.
Checklist
- Event schema with id, type, created time and version.
- Outbox so events match committed changes.
- Retries with backoff over days, then disable and notify.
- Per-endpoint isolation and timeouts.
- HMAC signatures with timestamps; SSRF protection.
- At-least-once, unordered delivery documented; dedupe by event id.
- Delivery logs, replay and an events API for reconciliation.