Building Inference Autopsy: Who Killed My TTFT?
I kept asking the same stupid question: who killed my TTFT? I would send a request to an OpenAI-compatible endpoint, get a plausible answer, and then watch the first token take forever as soon as the prompt or concurrency changed.
The total latency did not tell me enough. Sometimes the model was slow, but sometimes retrieval had stalled, the prompt had grown unexpectedly, a tool had retried, or the server was simply under pressure from another request.
That became the reason for building Inference Autopsy. I wanted a black-box profiler that could replay a workload and leave behind a trace instead of asking me to trust one number printed at the end.
Each trace follows the request through the parts I can observe: request setup, retrieval, prompt assembly, time to first token, inter-token gaps, tool calls, errors, and the final response. The point is not to pretend I can see inside every model; it is to make the boundary around the model less mysterious.
Concurrency made the problem more interesting. One request can look fine while a queue builds behind it, so I needed workload profiles, controlled limits, and enough evidence to see which request moved the tail.
I am also keeping deterministic fixtures beside live runs. Fixtures make regressions repeatable, while live endpoints show me the messy behaviour that a perfect local stub will never invent.
The project is still a profiler, not a magic diagnosis button. What it gives me is a better starting point: a timeline, a replayable workload, and a concrete question to investigate instead of the feeling that the LLM is just being slow today.