Uptime

MTTA

Also called mean time to acknowledge.

MTTA is the average time between an alert firing and a human acknowledging it. It measures how fast a problem gets an owner.

How it is measured

For each alert, subtract the fire timestamp from the acknowledge timestamp, then average across a period. Use the median and the 90th percentile alongside it, since one missed 3 a.m. page distorts a mean.

Segment by hour and severity. An MTTA of 4 minutes in the day and 22 at night points to rotation or notification settings, not alertness.

Worked example

A fintech's on-call tool logs 18 pages last month. Fourteen are acknowledged within 2 minutes. Three come at night and take 9, 14, and 31 minutes. One at 03:10 goes unacknowledged and is picked up by the secondary after 20. The median is 2 minutes, but the mean is 5.7.

The three slow ones all landed between 01:00 and 04:00, with the phone on silent. Adding a call-through for critical severity cuts the night median to 3 minutes.

How it differs

MTTA covers alert to acknowledgement. MTTR covers failure start to restored service. MTTA excludes the fix and the time before the alert. MTTR includes diagnosis and repair, and usually the detection gap too.

Common errors

Counting auto-acks from bots. Including muted test alerts. Ignoring percentiles. Starting the clock at human notice instead of alert time. Using it to blame individuals.

In practice

Pull acknowledgement times for the last 30 days, split day and night, and check the slowest five. Fix delivery, such as a phone call instead of a push, before you chase people. Turn off auto-acknowledge integrations.

See also

MTTR, MTTD

Sources

Count this on a real site.

Watch my website