My First Eval Set Was Too Nice
My first eval set was too nice to the system. Every example had the right fields, a clear request, and an answer that was easy to grade, so it mostly proved that the happy path worked.
Real messages are not like that. People leave out the important detail, change their mind halfway through a sentence, use a name the database does not know, or ask for two things when the workflow only has room for one.
I started adding those cases to the set: missing fields, ambiguous intent, wrong-but-plausible answers, and requests that should be refused instead of completed. The awkward examples were more useful than the polished ones because they showed me where the system was making assumptions.
I also learned that an eval is not just a score at the end. Each example should tell me what I am checking, what a good answer looks like, and what kind of failure I am trying to catch.
That made the eval set feel more like a small product specification. When I changed a prompt or a retrieval rule, I could see which behaviour moved instead of arguing with one overall number.
I am still adding examples as I find new failure modes. A useful eval set is not a trophy you finish; it is the record of the ways the system has surprised you so far.