Uptime
MTTD
Also called mean time to detect.
MTTD is the average time from when a problem starts to when someone or something detects it. It is the gap in which users suffer silently.
How it is measured
For each incident, subtract true start time from detection time. The true start comes from logs, the first failed request, or a retroactive probe, not from the alert. Average across incidents and look at the worst one.
It depends on probe interval and confirmation count: 5-minute checks with 3 confirmations can need 15 minutes. User reports as the first signal mean your monitoring is missing something.
Worked example
A donation page for a charity starts rejecting card tokens at 20:15 because a payment script URL changed. The HTTP check only loads the page, which stays 200. The first signal is a volunteer's email at 21:40.
MTTD for that event: 85 minutes. After adding a synthetic checkout step against the test gateway every 2 minutes with two confirmations, a replay of the failure would be caught in about 5.
How it differs
MTTD ends when the problem is detected. MTTA begins at detection and ends when a human takes it. MTTD excludes human response. MTTA excludes the unseen period before the alert.
Common errors
Taking alert time as start time. Averaging only incidents the monitors caught. Slow probe intervals. Relying on customer emails. Not backfilling the start from logs.
In practice
For the last five incidents, find the real start in logs and compare it to the first alert. For any gap over 10 minutes, add a probe or a heartbeat that would have caught it. Count customer-reported incidents separately.