FINANCE JUDGMENT CHALLENGE 001
THE PROVEN
RISK MODEL
A regional bank's AI credit-risk model has called this right for five years — 96% accuracy across 40,000 loans. Then rates jumped the fastest in a decade. The model still says the portfolio is low risk.
What would you do?
THE SITUATION
The record is real. The conditions behind it just changed.
A regional bank's AI credit-risk model has, over five years and roughly 40,000 loans, correctly classified low-risk accounts 96% of the time.
Three weeks ago, the central bank raised interest rates sharply and unexpectedly — the fastest move in a decade. Debt-service costs have jumped overnight across the bank's entire lending book.
The model hasn't been retrained. It still classifies the bank's mid-market manufacturing portfolio as low-risk, and a new $40 million lending tranche into that same portfolio is up for approval this week.
A competitor bank is reportedly already extending credit to some of the same clients. Waiting a full quarter for post-shock data risks losing those relationships.
Nothing about the model has visibly failed.
MAKE THE CALL
What do you do?
Choose before you continue.
WATCH THE CHALLENGE
Coming January 7.
The video for this challenge publishes January 7. Check back then — or work through the situation and questions below in the meantime.
THE JUDGMENT PROBLEM
Reliable then isn't the same as relevant now.
The model isn't wrong about what it learned. It's producing a highly reliable prediction, based on genuinely excellent historical performance, under conditions that just changed.
The question isn't whether the model has earned trust. Clearly it has. It's whether evidence gathered under one set of conditions transfers to a decision being made under another — and how much of it survives.
A 96% track record is evidence about how the model performed under the conditions represented in that record. The rate shock changed one of those conditions materially.
That doesn't prove the model is wrong. It means the bank has to judge how much its history still tells them about borrowers whose debt-service burden just changed.
THE EVIDENCE PROBLEM
A track record travels only as far as its conditions do.
The historical reliability didn't disappear. What changed is how much of it is still relevant evidence today.
AI can score the portfolio perfectly. Judgment determines how much of its history still applies.
THE JUDGMENT DIFFERENCE
Reliable is not the same as still relevant.
Low risk. 96% historically accurate.
Under today's conditions, is that still true?
Five years of evidence, three weeks of a new regime. How much travels?
BETTER QUESTIONS
Before you approve the tranche, ask:
Which borrowers carry the most rate exposure, and would a rough view change today's picture more than the model's score?
Is there a faster signal than waiting a full quarter — early repayment data, a covenant check, a call to the largest exposures?
What does each mistake cost: losing the relationship by waiting, or extending $40M into risk the model hasn't confirmed it still accounts for?
This isn't about trusting or distrusting the model. It's about how much of its evidence still applies right now.
SO, WHAT WOULD I DO?
Each option is defensible — depending on how much of the record still travels.
A trusts a real, earned track record — and treats a three-week-old shock as if it doesn't matter. B protects against committing capital before the changed risk is understood — but may cost the relationship. C hedges exposure without resolving whether the remaining exposure is appropriately priced. D keeps the relationship with a safeguard — but the safeguard's threshold still requires the judgment nobody's made yet.
This isn't about trusting or distrusting the model in general. It's about how much of its evidence still applies right now.
THE FINANCE JUDGMENT TAKEAWAY
A track record is evidence about the conditions it was earned under.
Ask what changed, not whether to trust the model in general.
Know how much of the evidence still travels before you act.
The historical reliability didn't disappear.
What changed is how much of it is still relevant today.
BRING THE CHALLENGE TO YOUR TEAM
Make judgment visible.
Judgment Challenges can be explored with your team through a private 90-minute Judgment Lab, turning the scenario into a practical conversation about model reliability, changing conditions, and how much history should count.