← All blogs

Your Average Is Lying to You: A Field Guide to P50, P95, and P99

Averages hide the users you are failing. How P50, P95, and P99 work, why the tail matters more than it looks, and the gotchas that corrupt APM dashboards.

DexoCode Team7 min read
  • APM
  • Performance
  • Observability
Your average is lying

Picture this. Your dashboard says the average response time for your checkout API is 150 milliseconds. Fast. Green. You lean back. Then support pings you: three customers are furious because checkout "took forever." You pull up the logs, and sure enough, some requests took almost four full seconds.

So which is it? Is your app fast or slow?

The honest answer is both, and the reason you can't see it is that you have been staring at the wrong number. The average is a smooth liar. It hides your worst moments inside a comfortable-looking middle. Percentiles are how you catch it lying, and P50, P95, and P99 are the three numbers that actually run a serious APM setup.

Let me show you why.

The average hides the people you are hurting

Here is a real-looking batch of 100 requests to that checkout endpoint. I have grouped them by how long they took:

  • 50 requests finished in about 40 ms
  • 45 requests finished in about 120 ms
  • 4 requests took around 900 ms
  • 1 request took a painful 4000 ms

Add it all up and divide by 100 and you get an average of 150 ms. That is the number on your dashboard. It looks fine.

Now watch what percentiles do to that same data.

A percentile answers a very specific question: "What was the slowest experience for the fastest X percent of requests?" Sort every request from quickest to slowest, then walk up the line.

  • P50 is the value in the middle. Half of requests were faster, half were slower. Here P50 is 40 ms. The typical user is having a great time.
  • P95 is the point where 95 percent of requests were faster. Here that is 120 ms. Still totally fine.
  • P99 is the point where 99 percent were faster, meaning the slowest 1 percent live beyond it. Here P99 is 900 ms. Almost a full second.

So the same endpoint, in the same window, is 40 ms for the median user and 900 ms for one in every hundred. The average of 150 ms describes literally nobody's actual experience. No single user got 150 ms. It is a mathematical ghost.

This is the whole point of percentiles. P50 tells you about your typical user. P95 and P99 tell you about the users you are quietly failing. If you only watch averages, those failing users are invisible until they complain, and by then they have already decided your product is flaky.

Why P99 matters more than it looks like it should

"One percent," you might say. "So what. Ninety-nine percent of requests are fast. Ship it."

That instinct feels reasonable and it is wrong, for two reasons that most people never think about.

Reason one: a single page is not a single request

Modern apps are chatty. Loading one screen might fire fifty, eighty, a hundred backend calls: the user profile, the cart, recommendations, pricing, inventory, ads, fonts, feature flags, and so on. Each of those calls rolls the dice against your latency distribution.

If each call independently has a 1 percent chance of landing in that slow tail, and a page makes 100 calls, then the chance the page escapes with zero slow calls is 0.99 to the power of 100, which is about 0.37. Flip that around. Roughly 63 percent of page loads will hit at least one slow call.

Read that again. Your "1 in 100" problem just became a "2 out of 3" problem at the page level. The tail is not a rare edge case. It is the default experience once you fan out across enough calls. This is why teams at scale obsess over P99 and even P99.9. Their tail becomes almost everyone's typical.

Reason two: your best customers hit the tail the most

The user who clicks around all day, loads a hundred screens, and generates thousands of requests is the one most likely to run into your slowest responses. The more someone uses your product, the more chances they have to feel your worst moment. In other words, your power users, the people you least want to annoy, are statistically guaranteed to meet your P99 over and over. Averages will never tell you this.

How this actually shows up in APM

Every real APM tool, whether it is Datadog, New Relic, Grafana, Honeycomb, or something you built in-house, gives you percentile views for a reason. Here is how to put them to work instead of just looking at them.

Watch the shape, not one number. Put P50, P95, and P99 on the same graph. The gap between them is the story. If P50 is flat and calm while P99 spikes, you have a tail problem: something occasional and nasty, like garbage collection pauses, a cold cache, a slow database query on certain inputs, or lock contention under load. If all three rise together, the whole system is under pressure, probably capacity or a downstream dependency. Two very different diagnoses, and you can only tell them apart by looking at the spread.

Alert on percentiles, not averages. An alert on average latency is close to useless because a handful of fast requests will paper over a growing tail. Alerting when P99 crosses a threshold catches real pain that real people are feeling. A common setup is a page or ticket when P99 breaches your target for several minutes straight.

Tie percentiles to money and to SLOs. This is where percentiles stop being a math curiosity and become a business tool. A Service Level Objective is usually written as a percentile promise, something like "99 percent of checkout requests complete in under 300 ms." That single sentence gives you an error budget. If you are allowed to miss on 1 percent of requests, you know exactly how much slowness you can spend before you have to stop shipping features and go fix performance. And the money link is real. The classic industry findings, from Amazon and Google onward, kept showing that small latency increases measurably drop conversions and engagement. The tail is not vanity. It is revenue leaking out slowly.

The gotchas that trip up almost everyone

This is the part most articles skip, so pay attention, because these mistakes quietly corrupt dashboards at real companies every day.

You cannot average percentiles. Ever. Say you have five servers, each reporting its own P99. It is extremely tempting to average those five numbers and call it your fleet P99. That answer is not just imprecise, it is meaningless. A percentile is a property of a distribution, and you cannot recover the combined distribution by averaging summaries of the pieces. Imagine four quiet servers and one server melting down. Its P99 is huge, the others are tiny, the average looks moderate, and your actual users on the bad server are on fire. To get a correct fleet-wide percentile, your APM has to merge the raw distributions, which is why good tools store latency as histograms or sketches like t-digest or HDR histograms rather than storing pre-computed percentiles. If your tooling shows you an "average of P99 across hosts," treat it with deep suspicion.

A percentile means nothing without a window. P99 over the last one minute and P99 over the last twenty-four hours are completely different animals. The daily number smooths over your 3 a.m. deploy that briefly made everything terrible. The one-minute number shows the pain as it happens but jumps around a lot. When someone quotes you a P99, always ask "over what time range," or you are comparing numbers that were never comparable.

Low traffic makes high percentiles jittery. P99 needs volume to mean anything. If an endpoint only got 50 requests in your window, its "P99" is basically defined by a single unlucky request, so it will bounce wildly for no real reason. On low-traffic services, P95 is often steadier and more honest than P99, and alerting on a shaky P99 there just trains your team to ignore alarms.

P99 still hides the truly catastrophic stuff. P99 tells you about the slowest 1 in 100. It says nothing about the slowest 1 in 1000, which is where timeouts, retries, and full request failures like to hide. At scale, that 1 in 1000 is thousands of angry humans a day. This is why big platforms track P99.9 and P99.99. Each extra nine is a smaller, angrier group of users, and the deeper you look, the uglier and more interesting the failures get.

The short version to remember

P50 is your typical user, the one you high-five in the standup. P95 and P99 are the users you are failing without knowing it, and thanks to how apps fan out across many calls, that tail reaches far more people than the raw percentage suggests. Averages will always tell you a comforting story. Percentiles tell you the true one.

So next time your dashboard glows a calm green average, do not relax. Go look at the P99. That is where your product is actually being judged.

Contact Us

Reach us by email, phone, WhatsApp, or social, we'll get back to you.

DexoCode

A leading software development company based in Calicut - India, partnering with top-tier clients from innovative start-ups to established enterprises.