I Keep Adding One More Metric
I kept adding another metric to Inference Autopsy because the first number never explained the failure. A slow request can be a slow first token, a gap between tokens, a retrieval stall, a tool retry, or a response that took too long to validate.
The temptation is to put every one of those numbers on the screen. It feels thorough for about five minutes, and then the trace becomes another thing I have to interpret before I can start debugging.
I am trying to keep the useful split: a few headline timings for orientation, then the request-level evidence underneath. TTFT and p99 tell me where to look; the timeline tells me what actually happened.
The same applies to quality metrics. A latency improvement is not automatically a win if the answer got worse, and a better eval score is not very comforting if the endpoint now times out under ordinary concurrency.
I started treating each metric as a question instead of a decoration. If I cannot say what decision the number helps me make, it probably does not belong in the first view.
I still add metrics sometimes. The difference is that I now try to remove one when I add one, so the profiler stays a debugging tool instead of quietly turning into a dashboard project.