Evaluation sets before prompts
If you cannot measure whether the assistant got better, you cannot ship it safely. Build the test set first.
By Amrelia Technologies Engineering · Engineering team
Every AI project we take on starts the same way: collect fifty to two hundred real questions with known good answers before writing a single prompt. Here is why, and how.
Why first
Without a test set, every change is judged by vibes. With one, you can swap models, change chunking or tighten retrieval and know within minutes whether quality moved.
How to build one
- Pull real questions from tickets, search logs and Slack.
- Have a domain expert write or approve the answer and the source.
- Tag each with difficulty and type: factual, procedural, multi-document, unanswerable.
- Include questions the assistant should refuse.
Run it on every change
In CI, like any other test. Track answer correctness, citation accuracy and refusal behaviour. If a number drops, the change does not ship.