All Tabs

I Started With a Stopwatch. That Was Not Enough.

aiengineeringobservabilitylessons

Inference Autopsy started as a stopwatch, but a number never told me why a run was slow. I kept the traces because they give me something to inspect instead of another average to argue about.

At first I measured the request from start to finish and called that latency. It was easy to print, but not very useful when the same total could come from a slow first token, a retrieval pause, or a tool that quietly retried.

The trace made me split the request into things I could reason about: time to first token, token gaps, retrieval, prompt assembly, tool calls, and the final answer. Suddenly “slow” was not one mystery anymore.

The other lesson was that averages are not enough. A run can look healthy in the middle while one unlucky request gets stuck in the tail, so I need the request-level evidence behind p95 and p99.

Replay mattered too. Once I could run the same workload again, I could tell whether a change improved the system or whether I had simply caught a quieter moment on the server.

I still like a simple stopwatch for a quick check. I just do not confuse a stopwatch reading with an explanation anymore; the explanation lives in the timeline around it.

Thanks for reading.Read more tabs →