Skip to content
All articles

Observability · Latency

Same average, very different users

Two services with an identical 153 ms mean latency: one where nobody waits more than 161 ms, one where a thousand requests take over a second. Why the mean cannot tell them apart.

· 4 min read

Two latency panels with the same 153 ms mean. Service A: p50 153 ms, p99 161 ms, no requests over 1 second. Service B: p50 14 ms, p99 1,416 ms, 1,000 of 10,000 requests over 1 second.

Two services each handle 10,000 requests. A latency dashboard that shows the mean reports the same number for both: 153.0 ms, identical to one decimal place. They are not the same service.

Meanp50p90p99MaxRequests over 1 s
Service A, uniformly a little slow153.0 ms153.0 ms160.0 ms161.0 ms161.0 ms0
Service B, fast with a 10% tail153.0 ms14.0 ms17.0 ms1,416.0 ms1,419.0 ms1,000

In service A, nobody waits more than 161 ms. In service B, a thousand people wait almost a second and a half.

The mean is doing exactly its job

A mean measures central tendency. It is deliberately insensitive to outliers, and in latency the outliers are the incident. Asking the mean to show a slow tail is asking it to do the one thing it was designed not to do.

Two things make it worse in practice. First, the mean has no unit of experience: no user of service B ever waited 153 ms. Their requests took about 14 ms or about 1,400 ms, and the average describes a request that never happened.

Second, the tail is not a small group of people. A p99 affects one request in a hundred, but a page or session that makes 100 requests is likely to hit it at least once: if each request independently has a 1% chance of landing in the tail, the chance is 1 − 0.99¹⁰⁰, about 63%. Fan-out turns a rare request into a common experience.

What to watch instead

  • Percentiles per endpoint: p50 for the typical request, p99 and the maximum for the people having a bad time.
  • A count of requests over a threshold you care about, such as one second. It reads like the table above, and it maps directly to an objective.
  • Histograms, aggregated across servers. They can be merged; percentiles cannot.

Common Mistake

Averaging each server's p99 to get a fleet-wide p99. Percentiles do not combine that way: the mean of eight servers' p99 values is not the p99 of all their requests. Merge the latency histograms first, then read the percentile from the result.

The mean still has a use: it tracks total work and capacity. It just cannot tell you who is waiting.

Get one diagram a week

A short article built around one engineering diagram, from the same library as these courses.

One diagram-led article a week on AI and systems engineering. We email you once to confirm, and every newsletter has an unsubscribe link. Privacy policy