Incident response and postmortems
How engineering teams handle production incidents: detection and paging, severity levels, incident roles, mitigation before diagnosis, communication with users, runbooks, blameless postmortems, action items, and designing systems that are easier to operate during an outage.
Reading is half of it. See this used in a real interview: walk through Design a Metrics and Monitoring System →
Every system fails eventually. What separates mature teams is how quickly they notice, how calmly they restore service, and whether the same failure happens again. Senior and staff interviews sometimes ask "tell me how you would handle an outage of this system" or "what happens at 3 a.m. when this breaks?". Knowing incident practice also improves designs: you build in the switches, dashboards and fallbacks that responders will need.
Detection
- Alerts based on user-facing symptoms (error rate, latency, SLO burn rate) rather than every internal cause. See SLIs, SLOs and error budgets.
- Synthetic checks that exercise critical journeys continuously.
- Customer reports routed to engineering quickly.
- Pages go to a clearly defined on-call engineer, with escalation if unacknowledged.
Severity levels
Define levels so everyone reacts proportionally, for example:
| Level | Meaning | Response |
|---|---|---|
| SEV1 | major outage or data loss affecting many users | all hands, incident commander, status page, executive updates |
| SEV2 | significant degradation or a major feature down | incident commander, status page |
| SEV3 | minor impact or a workaround exists | handled by the owning team in working hours |
Roles
For significant incidents:
- Incident commander: coordinates, decides, keeps the timeline; does not debug.
- Operations or technical lead(s): investigate and apply changes.
- Communications lead: updates the status page, support teams and stakeholders on a regular schedule.
- Scribe: records actions and findings for the postmortem.
Clear roles prevent the common failure of ten engineers debugging while nobody communicates or decides.
Mitigate first, diagnose later
The priority is restoring service, not finding the root cause:
- Roll back the latest deploy or configuration change (most incidents follow a change). See deployment strategies.
- Turn off the offending feature with a flag. See feature flags and A/B testing.
- Fail over to another zone or region. See active-active vs active-passive.
- Shed load or rate-limit an abusive client. See load shedding and backpressure.
- Scale up if capacity is the problem.
Systems designed with these levers (kill switches, quick rollbacks, failover, per-tenant limits) shorten incidents dramatically. Investigation continues after users are no longer affected.
Tools during an incident
- Dashboards per service showing the golden signals (traffic, errors, latency, saturation), with deploy and config change markers.
- Traces and logs searchable by request and tenant. See distributed tracing.
- Runbooks for known failure modes: symptoms, checks, and safe actions.
- A shared incident channel and document with the timeline.
Communication
Users and internal teams need honest, regular updates: what is affected, what is being done, when the next update will come. A status page hosted independently of the main infrastructure stays up when the product does not. Enterprise customers may have contractual notification requirements. See multi-tenancy.
Postmortems
After resolution, write a blameless postmortem:
- Summary and customer impact (duration, users affected, data impact).
- Timeline of detection, response and mitigation.
- Root causes and contributing factors, focusing on systems and processes rather than individuals ("the deploy pipeline allowed a config change without canarying", not "Alex pushed a bad config").
- What went well and what was lucky.
- Action items with owners and deadlines: fix the bug, add the missing alert, add a guardrail, update the runbook.
Track action items to completion; a postmortem without follow-through repeats. Share postmortems widely so other teams learn.
Learning over time
- Review incident trends: common causes (deploys, config, dependencies, capacity) point to the highest-value investments.
- Use error budgets to balance reliability work against features. See SLIs, SLOs and error budgets.
- Practise with game days and chaos experiments. See testing distributed systems.
Designing for operability
When presenting a design, mention the operational hooks: per-service SLO dashboards, kill switches for risky features, the ability to fail over or drain a region, per-tenant rate limits, safe rollbacks, and runbooks for the main failure modes. For Design a Payment System, add reconciliation alerts that catch silent money errors. For Design a Monitoring System, note that the monitoring system itself must be more reliable than what it monitors, and should alert when it cannot see data.
Checklist
- Symptom-based alerts and clear on-call ownership.
- Severity levels with matching responses.
- Incident commander, technical lead, communications and scribe roles.
- Mitigate first: rollback, flags, failover, load shedding, scaling.
- Dashboards with change markers, traces, runbooks.
- Independent status page and regular updates.
- Blameless postmortems with tracked action items.