Skip to content
AmreliaTechnologies
AI6 min read

Evaluation sets before prompts

If you cannot measure whether the assistant got better, you cannot ship it safely. Build the test set first.

By Amrelia Technologies Engineering · Engineering team

Every AI project we take on starts the same way: collect fifty to two hundred real questions with known good answers before writing a single prompt. Here is why, and how.

Why first

Without a test set, every change is judged by vibes. With one, you can swap models, change chunking or tighten retrieval and know within minutes whether quality moved.

How to build one

  • Pull real questions from tickets, search logs and Slack.
  • Have a domain expert write or approve the answer and the source.
  • Tag each with difficulty and type: factual, procedural, multi-document, unanswerable.
  • Include questions the assistant should refuse.

Run it on every change

In CI, like any other test. Track answer correctness, citation accuracy and refusal behaviour. If a number drops, the change does not ship.

  • ai
  • rag
  • evaluation
LinkedInX

Related articles

Working on this right now?

We would rather talk through a real problem than write another article about it.