A test for the AI part

Normal software has tests: given this input, expect that output. The AI part of the software needs the same thing, with one difference. A model can give a slightly different answer each time, and a change that fixes one case can break another. An evaluation set catches that before your users do.

What goes in it

Part Example for an invoice workflow
The input An invoice PDF from your inbox
The right answer Supplier, invoice number, date, total, VAT, purchase order
The rule behind it “Totals in a foreign currency are converted at the invoice date”
A label Normal case, hard case, or case that must go to a person

The hard cases matter most: blurry scans, handwritten notes, two invoices in one file, a supplier with a new layout. A set with only easy cases gives a high score and a false sense of safety.

How it is used

The set runs on every change: a new prompt, a different model, a code change around the model. If a change makes the score worse, it does not ship. When the system makes a mistake in production, that case is added to the set, so the same mistake is caught next time.

Who owns it

You do. The evaluation set lives in your repository with the rest of the code. If you change suppliers, models or teams, the set still tells the next person exactly how good the system is on your cases.

How this applies to us

Every sprint starts the evaluation set on day one, from 100 to 300 of your real cases, before the AI part is finished. In a proof of concept, the evaluation score is the main result: it tells you whether the idea is accurate enough to build for production.