News
Two numbers, never one: measuring an agent like a fraud rule
One accuracy number hides the thing that will get your project cancelled.

Production notes
Frontal Designs


Accurate at what?
Almost every agent project we are asked to look at reports a single figure. It is 94% accurate. If that missing 6% is 1% real problems missed and 5% good cases bounced to a human for no reason, you have built a machine that creates five times more work than it removes. Your ops team will stop using it within a quarter and they will not tell you why.
Miss rate: real problems the agent let through.
Over-flag rate: good cases it wrongly stopped or escalated.
The two carry different prices. Always.
How fraud teams report it
A fraud rule ships with both numbers, measured against a labelled holdout set, and a budget for each agreed with the business side. Breach the budget and someone with a name is accountable. We report agents exactly the same way, weekly, and the process owner sees both lines.
Setting the thresholds
Start from cost. What does a wrong yes cost in money? What does a wrong no cost? For invoice posting a wrong yes is a mis-posting that is caught at month end. A wrong no is a supplier paid late. For KYC the prices flip. The thresholds follow the prices, not a round number.
What changes in the build
The confidence gate becomes three-way instead of two. The exception queue gets a reason code so you can see which error type is growing. The weekly report has two lines and a trend. None of this is expensive. All of it is skipped by teams measuring one number.
If your dashboard shows one accuracy figure, ask for the other one. The answer will tell you a lot.

Blog & Insight
Read More Notes
89% of AI agent pilots never reach production. What the other 11% did differently.





