A test for the AI part
Normal software has tests: given this input, expect that output. The AI part of the software needs the same thing, with one difference. A model can give a slightly different answer each time, and a change that fixes one case can break another. An evaluation set catches that before your users do.
What goes in it
| Part | Example for an invoice workflow |
|---|---|
| The input | An invoice PDF from your inbox |
| The right answer | Supplier, invoice number, date, total, VAT, purchase order |
| The rule behind it | “Totals in a foreign currency are converted at the invoice date” |
| A label | Normal case, hard case, or case that must go to a person |
The hard cases matter most: blurry scans, handwritten notes, two invoices in one file, a supplier with a new layout. A set with only easy cases gives a high score and a false sense of safety.
How it is used
The set runs on every change: a new prompt, a different model, a code change around the model. If a change makes the score worse, it does not ship. When the system makes a mistake in production, that case is added to the set, so the same mistake is caught next time.
Who owns it
You do. The evaluation set lives in your repository with the rest of the code. If you change suppliers, models or teams, the set still tells the next person exactly how good the system is on your cases.
How this applies to us
Every sprint starts the evaluation set on day one, from 100 to 300 of your real cases, before the AI part is finished. In a proof of concept, the evaluation score is the main result: it tells you whether the idea is accurate enough to build for production.