Server

p99 latency

p99 latency is the response time that 99 percent of requests beat. It exposes the slow tail that p50 and p95 smooth over.

How it is measured

Use a histogram with enough resolution at the high end, and a window long enough to hold hundreds of samples above the 99th. With 1,000 requests there are only 10 above it, so one outlier moves it.

Measure at more than one point if you can: the app, the proxy and the browser. The same p99 at all three means the delay is in the app, a gap at the proxy means queueing.

Worked example

A search API handles 600 requests per second. p95 is 220 ms, and p99 is 2.4 seconds. A trace of the slowest 1 percent shows each one waited about 2 seconds for a database connection because the pool of 15 was exhausted during bursts.

At that rate 6 requests every second, around 21,000 an hour, were slow. Raising the pool to 40 and fixing a long transaction brought p99 to 380 ms.

How it differs

p99 latency covers the slowest 1 percent. Tail latency is the broader idea of the slow end of the distribution, with p99 as one way to quantify it. A p99 excludes the very worst requests beyond it, such as p99.9. Tail latency excludes any single named percentile, and covers whatever your users feel most.

Common errors

Reading p99 from too few samples. Averaging p99 across hosts. Dropping slow requests as outliers. Believing p99 does not matter because it is only 1 percent, when a page that makes 50 calls will hit the tail on most views. Alerting on single-minute spikes.

In practice

Put p99 for your busiest endpoint on a dashboard, and match its spikes to GC logs, pool waits, cold starts and backups. When you find one cause, fix it and see whether the next one is hiding behind it.

See also

p95 latency, Tail latency

Sources

Count this on a real site.

Watch my website