All Tabs

p99 Ruined My Very Good Average

ailatencyinference-autopsy

I used to look at average latency and feel informed. Then I started building Inference Autopsy and p99 began following me around like a debt collector.

Imagine most requests finishing in two seconds while one stretches toward twenty. The average can still look respectable, but the person who got that request experiences a broken app, not a respectable number.

AI systems make the tail harder to explain because “slow” can mean several things. The first token might be late, tokens might arrive with gaps, retrieval might stall, or a tool might quietly retry before the model finishes.

That is why I keep request-level traces. When p99 moves, I want to find the requests in that tail, inspect their stages, and replay something close to the same workload instead of staring at the percentile by itself.

I still report averages. They are useful for knowing what normal felt like, but they should not speak for the people who had the bad day.

p50 tells me the middle, p95 and p99 tell me how rough the tail became, and the traces tell me who ruined it. That combination is much more actionable than any one number pretending to be the whole story.

Thanks for reading.Read more tabs →