News

Evaluation sets are the product, not the model

Two to five hundred labelled real cases beat a better model. Every time.

Author Image

Production notes

Frontal Designs

Image

The asset nobody wants to build

Collecting real cases with known correct answers is tedious. Two to five hundred invoices, KYC files or reconciliation breaks, each labelled by someone who knows the process. It takes weeks. Nobody gets thanked for it. It is also the reason the agent is still running a year later.

  • Agents with full evaluation coverage: 9% rollback rate in the last twelve months.

  • Agents without: 47%.

  • That is the difference between a system and an experiment.

What the set is for

A fixed set of real cases, a pass threshold agreed with the business, and a gate that blocks deployment when the number drops. Every prompt change, model swap or integration fix runs against it before it ships. At PayPal we never released a fraud rule without a holdout set. An agent is a fraud rule with more freedom.


Model upgrades are free once you have the set. Without it, every upgrade is a gamble.

Model upgrades are free once you have the set. Without it, every upgrade is a gamble.

Model upgrades are free once you have the set. Without it, every upgrade is a gamble.

Building it well

Sample from real traffic, not from the happy path. Include the photographed invoices, the supplier who changed their template, the KYC file with a smudged date. Label the two error types separately so you can report a miss rate and an over-flag rate rather than one accuracy number that hides both.

Keeping it alive

Cases from the exception queue feed back into the set each month. The set grows with the process. When the business changes a rule, the set changes first and the agent follows.


This is also where the cost argument is won. A team that can prove its agent against real cases can switch to a cheaper model tier with confidence. A team that cannot is stuck paying for the expensive one forever.
This is also where the cost argument is won. A team that can prove its agent against real cases can switch to a cheaper model tier with confidence. A team that cannot is stuck paying for the expensive one forever.
This is also where the cost argument is won. A team that can prove its agent against real cases can switch to a cheaper model tier with confidence. A team that cannot is stuck paying for the expensive one forever.

If you are choosing between a better model and a better evaluation set, buy the set.

Frequently asked questions

Questions We Get Asked

Straight answers on what we build, how it is measured, and what it costs.

What kind of agents do you build?

Task-focused agents for document and data workflows, reconciliation, onboarding checks and back-office operations. Each one owns a single process end to end and routes uncertain cases to a person.

How is this different from a chatbot or RPA?

Where does it run?

Where are you based?

How do I get started?

Which industries do you work in?

Frequently asked questions

Questions We Get Asked

Straight answers on what we build, how it is measured, and what it costs.

What kind of agents do you build?

Task-focused agents for document and data workflows, reconciliation, onboarding checks and back-office operations. Each one owns a single process end to end and routes uncertain cases to a person.

How is this different from a chatbot or RPA?

Where does it run?

Where are you based?

How do I get started?

Which industries do you work in?

Frequently asked questions

Questions We Get Asked

Straight answers on what we build, how it is measured, and what it costs.

What kind of agents do you build?

Task-focused agents for document and data workflows, reconciliation, onboarding checks and back-office operations. Each one owns a single process end to end and routes uncertain cases to a person.

How is this different from a chatbot or RPA?

Where does it run?

Where are you based?

How do I get started?

Which industries do you work in?

AI AGENTS THAT SURVIVE PRODUCTION

Got a Process an Agent Could Own?

Tell us what it is. We will tell you straight whether it is worth building.

AI AGENTS THAT SURVIVE PRODUCTION

Got a Process an Agent Could Own?

Tell us what it is. We will tell you straight whether it is worth building.

AI AGENTS THAT SURVIVE PRODUCTION

Got a Process an Agent Could Own?

Tell us what it is. We will tell you straight whether it is worth building.