News
Evaluation sets are the product, not the model
Two to five hundred labelled real cases beat a better model. Every time.

Production notes
Frontal Designs


The asset nobody wants to build
Collecting real cases with known correct answers is tedious. Two to five hundred invoices, KYC files or reconciliation breaks, each labelled by someone who knows the process. It takes weeks. Nobody gets thanked for it. It is also the reason the agent is still running a year later.
Agents with full evaluation coverage: 9% rollback rate in the last twelve months.
Agents without: 47%.
That is the difference between a system and an experiment.
What the set is for
A fixed set of real cases, a pass threshold agreed with the business, and a gate that blocks deployment when the number drops. Every prompt change, model swap or integration fix runs against it before it ships. At PayPal we never released a fraud rule without a holdout set. An agent is a fraud rule with more freedom.
Building it well
Sample from real traffic, not from the happy path. Include the photographed invoices, the supplier who changed their template, the KYC file with a smudged date. Label the two error types separately so you can report a miss rate and an over-flag rate rather than one accuracy number that hides both.
Keeping it alive
Cases from the exception queue feed back into the set each month. The set grows with the process. When the business changes a rule, the set changes first and the agent follows.
If you are choosing between a better model and a better evaluation set, buy the set.

Blog & Insight
Read More Notes
89% of AI agent pilots never reach production. What the other 11% did differently.





