Server
p95 latency
p95 latency is the response time that 95 percent of requests beat. One in twenty is slower than this figure.
How it is measured
Take the 95th percentile of request durations from a histogram over a window of at least a few minutes. With 200 requests in the window, the 95th is only the 10th-slowest request, so use longer windows for low-traffic endpoints.
Break it down by route and by cache status. A p95 that mixes cached 20 ms pages and uncached 2 second pages tells you little.
Worked example
A checkout endpoint serves 40,000 requests a day with p50 of 190 ms and p95 of 1.6 seconds. Looking at slow traces, the long ones all call the shipping-rate API, which answers in 1.2 to 1.4 seconds when its own cache is cold.
Caching shipping rates per postcode for 15 minutes lowers p95 to 420 ms. The median barely changes, which is why p50 would never have shown it.
How it differs
p95 latency marks the line where the slowest 5 percent begin. p50 latency marks the middle. p99 latency marks the slowest 1 percent. p95 excludes the extreme tail and is more stable than p99 on small samples. It excludes the typical case that p50 describes, so use all three.
Common errors
Computing p95 from per-minute averages. Using a window so short the number jumps around. Excluding errors and timeouts from the sample. Setting an SLO on p95 and ignoring the tail. Comparing p95 for a cached and an uncached route in the same chart.
In practice
Choose a p95 target for each critical endpoint, such as 500 ms for checkout, and alert when it is exceeded for 10 minutes. When it breaks, open a trace of a slow request, not an average.