Server
Tail latency
Tail latency is the slow end of the response-time distribution, the requests far worse than the typical one. Users remember the tail because they hit it more often than the percentages suggest.
How it is measured
Examine the high percentiles (p99, p99.9) and the maximum, from histograms, not averages. Also look at the shape: a long thin tail comes from rare events like GC or a lock, and a fat tail comes from queueing under load.
Trace the slowest 1 percent and group by cause: a cold cache, a pool wait, a retry, a noisy neighbour. One or two causes usually explain most of them.
Worked example
A product page makes 30 backend calls in parallel. Each call has p99 of 300 ms and p50 of 20 ms. The chance that at least one of 30 calls lands in the slow 1 percent is about 26 percent, so roughly a quarter of page loads feel slow.
Adding a 50 ms hedged retry on the slowest call, with caching on two others, moved page p95 from 310 ms to 90 ms. No individual call got faster at the median.
How it differs
Tail latency is the concept: the bad end of the distribution. p99 latency is one number that measures it. Tail latency excludes the typical case and the mean, and a p99 excludes extremes beyond the 99th percentile. A system can have a fast p50 and a miserable tail.
Common errors
Reporting averages. Ignoring fan-out, where many calls make the tail likely. Discarding slow requests as noise. Fixing the median when the tail is the complaint. Setting timeouts so long that the tail has no ceiling.
In practice
Trace the 20 slowest requests from yesterday and list what each was waiting on. Fix the most common cause first, and consider timeouts, hedged requests or caching on high fan-out paths.