The Model Was Not the Whole System
I started by measuring the model because it was the obvious part. Then retrieval, prompt assembly, tools, and answer quality kept breaking around it, so now I treat the whole system as the thing being profiled.
A request can wait for a document before the model sees anything, spend time building a prompt, call a tool twice, and still finish with a fast generation phase. Looking only at model latency would make that request look healthy when the user experienced the whole wait.
I began adding stages to the trace so each part could be inspected on its own. Retrieval needs its query and result set, prompt assembly needs its size and timing, and tool calls need inputs, outcomes, and the reason they happened.
Answer quality belongs in the same conversation. A fast answer that misses the retrieved fact is not a performance win, and a slower answer may be acceptable if the extra time came from a deliberate verification step.
This changed the way I think about regressions. The question is no longer just “did the model get slower?” but “which stage changed, under which workload, and did the user-visible result get better or worse?”
The profiler is becoming a system tool because AI products are systems. The model is important, but the surrounding data, tools, policies, and evaluation path decide whether the model can do anything useful with the request.