All Tabs

What I Read Before Building the Profiler

aiengineeringinferencelearning

I made a reading list before building the profiler because I did not want to invent my own definitions for TTFT, ITL, or a “good” benchmark. The list is less exciting than writing code, but it keeps me from building a very confident stopwatch with the wrong maths.

I started with benchmarking and serving behaviour: batching, queueing, streaming, and the difference between a request that is slow once and a system that is slow under load. Those details decide what a latency number means before my code ever records it.

Then I went deeper into tracing and async work. I wanted to understand how to preserve ordering, record partial failures, and keep a streamed response useful even when one stage does not behave perfectly.

Evaluation was another part of the preparation. A profiler that only measures speed can celebrate a faster system that gives worse answers, so the benchmark needs a way to connect runtime evidence with answer quality.

I also looked at CLI design and the shape of reports. If the output is too dense, people will skip it; if it hides the raw evidence, they cannot check the conclusion. The interface has to make the investigation easier, not turn the result into a second mystery.

The reading did not give me a perfect architecture. It gave me better questions and fewer accidental assumptions, which was exactly what I needed before I started writing the parts that would be hardest to change later.

Thanks for reading.Read more tabs →