FINANCE JUDGMENT CHALLENGE 002
THE FRAUD
THRESHOLD
A bank's AI fraud system scores every transaction 0–100. Lowering the freeze threshold from 90 catches millions more in fraud — and freezes thousands more legitimate payments, disproportionately from small suppliers.
What would you do?
THE SITUATION
The model can price the trade-off precisely. It can't decide it.
A bank's AI fraud system scores every outgoing transaction from 0 to 100. For the past year, transactions scoring 90 or above have been automatically frozen pending manual review.
The Head of Fraud Operations brings the Chief Risk Officer a proposal. Lowering the threshold to 82 would have caught an additional €4.2 million in fraud last year. At 86, roughly €2.5 million — with substantially fewer false positives than at 82.
The false positives hit both small and large accounts, but the consequences aren't equal. A large corporate can get a frozen payment reviewed within hours. For a small supplier, the same false positive can mean a payroll run sits frozen for two or three days.
Same model error. Unequal consequence — and it grows sharply below 86.
MAKE THE CALL
What do you do?
Choose before you continue.
WATCH THE CHALLENGE
Coming January 14.
The video for this challenge publishes January 14. Check back then — or work through the situation and questions below in the meantime.
THE JUDGMENT PROBLEM
Precise isn't the same as decided.
The model has done its job. It can show, with real precision, what each threshold catches and what each threshold costs.
The question isn't whether that trade-off calculation is trustworthy. It's who decides what a percentage point of fraud prevention is worth, in disrupted legitimate business — and who decides which customers absorb that disruption.
A threshold isn't a fact the model discovers. It's a value judgment the model has been asked to enforce, once someone upstream decides where to set it.
The bank largely bears the cost of undetected fraud. Customers bear the cost of false positives — and not evenly among themselves.
THE COST PROBLEM
One optimization, two kinds of cost.
The bank bears one. Customers bear the other — and not evenly.
AI can quantify the trade-off perfectly. Judgment determines what it's worth, and who bears it.
THE JUDGMENT DIFFERENCE
Performing well is not the same as being fair.
How much fraud is caught, how many payments freeze.
Who absorbs that freeze, and is that acceptable?
Same model, different value call at every threshold.
BETTER QUESTIONS
Before you move the threshold, ask:
Whose transactions get disrupted at each threshold — has anyone actually looked?
Is there a cheaper way to fix a false positive itself, rather than only moving where the line sits?
If this threshold is wrong, who bears that cost first — the bank, in fraud losses, or the customer, in a frozen payroll run?
This isn't about which threshold performs best. It's about who is actually paying for that performance.
SO, WHAT WOULD I DO?
None of these is free — and none is obviously the responsible one.
A keeps a known process — and leaves identified fraud on the table. B captures the full benefit — at the point where the burden falls hardest on small suppliers. C targets the disparity directly — but is still a judgment call dressed as a number. D looks sophisticated but multiplies the same question across categories instead of answering it.
This isn't about which threshold performs best. It's about who is actually paying for that performance.
THE FINANCE JUDGMENT TAKEAWAY
AI can quantify a trade-off. It can't decide what it's worth.
Ask who bears each kind of cost, not just which number performs best.
A technical setting is a standing decision about whose cost you accept.
AI can quantify a trade-off.
It can't decide what it's worth — or who should pay for it.
BRING THE CHALLENGE TO YOUR TEAM
Make judgment visible.
Judgment Challenges can be explored with your team through a private 90-minute Judgment Lab, turning the scenario into a practical conversation about fraud thresholds, disparate impact, and who a trade-off is actually asking to pay.