Measure the Tail, Not Just the Average

An average compresses a distribution into one comfortable number. Production latency is rarely comfortable or evenly distributed.

Suppose 99 requests finish in 40 milliseconds and one request takes four seconds. The average is about 80 milliseconds. That number describes almost nobody: most users saw half of it, while the unlucky user waited fifty times longer.

Keep the distribution

Track at least a few percentiles and request volume together:

MetricValue
P5040 ms
P9558 ms
P994.0 s
Requests100

Percentiles also need enough samples. A P99 calculated from a tiny window is mostly a story about one request, so retain histograms and compare equivalent traffic windows.

Find the population behind the tail

Split the slow requests by endpoint, region, payload size, cache state, dependency, and retry count. Tail latency often belongs to a specific population that disappears when everything is aggregated.

Optimizing the average rewards the common path. Reliability work begins when you ask who is still waiting.