SysDesignPrep.com
Study guide 179 of 183

Incident response and postmortems

How engineering teams handle production incidents: detection and paging, severity levels, incident roles, mitigation before diagnosis, communication with users, runbooks, blameless postmortems, action items, and designing systems that are easier to operate during an outage.

Reading is half of it. See this used in a real interview: walk through Design a Metrics and Monitoring System →

Every system fails eventually. What separates mature teams is how quickly they notice, how calmly they restore service, and whether the same failure happens again. Senior and staff interviews sometimes ask "tell me how you would handle an outage of this system" or "what happens at 3 a.m. when this breaks?". Knowing incident practice also improves designs: you build in the switches, dashboards and fallbacks that responders will need.

Detection

  • Alerts based on user-facing symptoms (error rate, latency, SLO burn rate) rather than every internal cause. See SLIs, SLOs and error budgets.
  • Synthetic checks that exercise critical journeys continuously.
  • Customer reports routed to engineering quickly.
  • Pages go to a clearly defined on-call engineer, with escalation if unacknowledged.

Severity levels

Define levels so everyone reacts proportionally, for example:

LevelMeaningResponse
SEV1major outage or data loss affecting many usersall hands, incident commander, status page, executive updates
SEV2significant degradation or a major feature downincident commander, status page
SEV3minor impact or a workaround existshandled by the owning team in working hours

Roles

For significant incidents:

  • Incident commander: coordinates, decides, keeps the timeline; does not debug.
  • Operations or technical lead(s): investigate and apply changes.
  • Communications lead: updates the status page, support teams and stakeholders on a regular schedule.
  • Scribe: records actions and findings for the postmortem.

Clear roles prevent the common failure of ten engineers debugging while nobody communicates or decides.

Mitigate first, diagnose later

The priority is restoring service, not finding the root cause:

Systems designed with these levers (kill switches, quick rollbacks, failover, per-tenant limits) shorten incidents dramatically. Investigation continues after users are no longer affected.

Tools during an incident

  • Dashboards per service showing the golden signals (traffic, errors, latency, saturation), with deploy and config change markers.
  • Traces and logs searchable by request and tenant. See distributed tracing.
  • Runbooks for known failure modes: symptoms, checks, and safe actions.
  • A shared incident channel and document with the timeline.

Communication

Users and internal teams need honest, regular updates: what is affected, what is being done, when the next update will come. A status page hosted independently of the main infrastructure stays up when the product does not. Enterprise customers may have contractual notification requirements. See multi-tenancy.

Postmortems

After resolution, write a blameless postmortem:

  • Summary and customer impact (duration, users affected, data impact).
  • Timeline of detection, response and mitigation.
  • Root causes and contributing factors, focusing on systems and processes rather than individuals ("the deploy pipeline allowed a config change without canarying", not "Alex pushed a bad config").
  • What went well and what was lucky.
  • Action items with owners and deadlines: fix the bug, add the missing alert, add a guardrail, update the runbook.

Track action items to completion; a postmortem without follow-through repeats. Share postmortems widely so other teams learn.

Learning over time

  • Review incident trends: common causes (deploys, config, dependencies, capacity) point to the highest-value investments.
  • Use error budgets to balance reliability work against features. See SLIs, SLOs and error budgets.
  • Practise with game days and chaos experiments. See testing distributed systems.

Designing for operability

When presenting a design, mention the operational hooks: per-service SLO dashboards, kill switches for risky features, the ability to fail over or drain a region, per-tenant rate limits, safe rollbacks, and runbooks for the main failure modes. For Design a Payment System, add reconciliation alerts that catch silent money errors. For Design a Monitoring System, note that the monitoring system itself must be more reliable than what it monitors, and should alert when it cannot see data.

Checklist

  • Symptom-based alerts and clear on-call ownership.
  • Severity levels with matching responses.
  • Incident commander, technical lead, communications and scribe roles.
  • Mitigate first: rollback, flags, failover, load shedding, scaling.
  • Dashboards with change markers, traces, runbooks.
  • Independent status page and regular updates.
  • Blameless postmortems with tracked action items.

Open in your browser to sign in

Google does not allow sign-in inside this app's built-in browser. Open this page in Safari and sign in there. The link opens this same page.

Tap the ⋯ or share button at the top or bottom of the screen, then Open in browser. Or copy the link and paste it into Safari.