All Tabs

One Request Was Not a Benchmark

aiengineeringprofilingbenchmarks

One request is a demo, not a benchmark. It can show that the endpoint answers, but it cannot tell me how the system behaves when prompts vary, streams overlap, or one request gets unlucky.

I started adding workload profiles so I could describe what I was actually running: prompt sizes, streaming, tool use, expected output, and the amount of concurrency. That was more useful than another flag called “stress test” with no clear meaning.

Replay was the part that changed my confidence. If I could run the same requests again, I could compare a code change against something stable instead of comparing two unrelated moments on a busy server.

Concurrency also needed to be controlled rather than shouted at the endpoint. A closed-loop runner can start another request when one finishes, which gives me a clearer picture of how the system behaves as pressure increases.

The tail is where the interesting bugs kept hiding. Median latency could improve while p95 got worse, and the only way to understand that was to find the individual requests and inspect their stages.

The next step is not to add more charts. It is to make the benchmark hard to misunderstand, with clear workload definitions, comparable traces, and enough context that another person can reproduce what I saw.

Thanks for reading.Read more tabs →