Uptime
Postmortem
Also called incident review.
Postmortem is a written review of an incident: what happened, why, how it was handled, and what will change. It is for learning, not blame.
How it is measured
It is not a metric, so confirm it exists and is useful. It should be done within five working days, with a timeline from logs, impact in minutes, users, or revenue, plural root causes, and action items with an owner and a due date.
Track follow-through: the percent of action items closed within 30 days. A postmortem that lists ten actions and closes none is a diary.
Worked example
A 52-minute outage at a bike-parts store came from a Redis eviction setting that dropped cart sessions after a 12 GB import. The writeup lists the timeline (13:05 import, 13:21 first failed carts, 13:57 maxmemory changed), the impact (310 abandoned carts, about 14,000 in lost orders), and three actions.
Action one, an alert on evicted keys, ships in two days. Action three, be more careful with imports, is rewritten as a staging dry run with a memory check.
How it differs
Postmortem is the written review after. Incident is the live event. The incident record holds the timeline and decisions made in the moment. The postmortem adds cause analysis and changes.
Common errors
Naming a person as the cause. Stopping at one root cause. Vague actions like improve monitoring. Writing it three weeks later. No owner on items. Keeping it private so no one else learns.
In practice
Hold one within a week of your next incident of any size. Use a template with timeline, impact, contributing factors, and actions. Put the actions in your tracker with dates and review them at the next team meeting.