SALES JUDGMENT CHALLENGE 007
THE
EXCEPTION
An AI opportunity-scoring system has closed 92% of its high-probability calls correctly for two years. This quarter, a €2M renewal scores 91 — and a 15-year rep says something's wrong, based on a pattern the model was never built to isolate.
What would you do?
THE SITUATION
A 92% track record. One deal that doesn't feel right.
An AI opportunity-scoring system rates every open deal on likelihood to close. Over two years and thousands of deals, its high-probability calls have closed correctly 92% of the time.
This quarter, a major renewal — €2 million — scores 91/100. The system has seen everything: response times, deal velocity, stakeholder engagement, including two specific changes — the champion's replies have gotten shorter, and he's stopped looping in his boss on calls the way he used to. Individually, neither has ever been strongly associated with lost deals. The score holds at 91.
The rep — fifteen years on the job — isn't convinced. She's seen this exact combination before, not either signal alone, and it's usually the last thing she notices before a deal quietly dies.
She can name exactly what she saw. She can't prove what it means.
MAKE THE CALL
What do you do?
Choose before you continue.
WATCH THE CHALLENGE
Coming November 23.
The video for this challenge publishes November 23. Check back then — or work through the situation and questions below in the meantime.
THE JUDGMENT PROBLEM
A track record is evidence. It isn't complete evidence.
The system didn't miss anything. It saw both changes and correctly noted that, on their own, neither predicts much.
The rep isn't claiming the AI is blind. She's claiming the combination, in this context, means something the model's history doesn't capture.
Nobody has proof — not the system, and not her. The question isn't "who's right?" It's: what should justify departing from a rule that's usually reliable?
THE EXCEPTION PROBLEM
92% tells you about the average deal. Not this one.
It doesn't settle whether this particular deal belongs to the minority of high-probability deals that still fail.
AI can calculate the odds perfectly. Judgment determines how much this case differs from the average — and what it costs to be wrong either way.
THE JUDGMENT DIFFERENCE
Reliable on average is not certain here.
91 out of 100. High probability.
Is this deal the exception?
The record can't see her pattern. Only testing it, inside the deal, can find out if it holds.
BETTER QUESTIONS
Before you act on either read, ask:
Has this rep's read been right before, on the record — not just this once?
Is there a real action inside the deal — not a data query — that would tell you more before you decide?
What does each kind of error cost: quietly losing a €2M renewal, or signaling concern where none may exist?
This isn't about whether to trust the model or the rep. It's about what would justify overriding either one — and what a wrong guess costs.
SO, WHAT WOULD I DO?
Each option is defensible — depending on what's at stake if I'm wrong.
A preserves forecasting discipline — and risks discounting a case-specific pattern that may matter more here than the model's historical weighting suggests. B honors real expertise, but a downgrade made before testing anything is indistinguishable from a confident guess. C tests the concern where it actually lives, but asking a champion to recommit can itself introduce doubt that wasn't there. D protects the relationship at the executive level, but escalating a healthy renewal can look like the company is worried when the client wasn't.
This isn't about who to trust in general. It's about what would justify overriding either one, in this specific case, given what's actually at stake.
THE SALES JUDGMENT TAKEAWAY
A strong track record deserves weight. So does a case it wasn't built to see.
Ask what would justify overriding either one.
Weigh what each kind of mistake would actually cost.
The question isn't whether to trust the AI or the rep.
It's what would justify overriding either one.
BRING THE CHALLENGE TO YOUR TEAM
Make judgment visible.
Judgment Challenges can be explored with your team through a private 90-minute Judgment Lab, turning the scenario into a practical conversation about forecasting discipline, instinct, and when a track record deserves to be overridden.