SALES JUDGMENT CHALLENGE 008
THE
BOUNDARY
An AI deal-approval agent has a 98% clean-outcome rate on renewals up to $50K. Leadership wants to triple its authority to $150K — but every one of those approvals was a smaller deal, and a mistake at the new tier costs three times more.
What would you do?
THE SITUATION
The track record is real. It was never measured at this tier.
An AI deal-approval agent is authorized to independently approve standard renewal contracts up to $50K in annual value — no human sign-off required. Over 18 months and 2,200 approvals, 98% completed cleanly: no disputes, no clawbacks, no compliance flags.
Renewals between $50K and $150K still require human review, and that queue now averages six business days. Of the deals sitting in that queue, 12% have seen the customer request an extension, delay signing, or start engaging a competitor during the wait.
Every one of the 2,200 approvals behind the 98% figure was a deal under $50K. Deals above $50K also tend to carry more negotiated terms, though nobody has established whether that changes how the agent decides.
Even if the agent stayed exactly as reliable at $150K, a mistake there costs three times more.
MAKE THE CALL
What do you do?
Choose before you continue.
WATCH THE CHALLENGE
Coming December 14.
The video for this challenge publishes December 14. Check back then — or work through the situation and questions below in the meantime.
THE JUDGMENT PROBLEM
Reliability and authority are not the same question.
This isn't "has the AI earned more trust?" On the numbers, clearly yes.
The harder question a strong track record can't answer by itself is: even if the AI stayed exactly as reliable at a higher tier, does more reliability automatically justify more authority — when what's at stake per mistake has also changed?
A 98% clean-outcome rate says something about how often the agent gets it right. It says nothing about what should happen with that information once the cost of the 2% changes.
A system can become more reliable without becoming more appropriate to authorize at every level of consequence. Those are two separate judgments.
THE AUTHORITY PROBLEM
A performance record can inform the decision. It can't make it for you.
Even a perfectly transferred 98% wouldn't settle whether $150K of unsupervised exposure is a boundary anyone should accept.
AI can hit 98% consistently. Judgment determines whether 98% is enough at this level of exposure.
THE JUDGMENT DIFFERENCE
Reliable is not the same as authorized.
98% clean, under $50K.
What does being wrong cost at $150K?
The rate might transfer perfectly. The stakes already have.
BETTER QUESTIONS
Before you raise the ceiling, ask:
Would you accept this same rate at $150K if it had been measured there directly?
Do the negotiated terms common above $50K actually change what the agent is deciding?
Is there a way to test the agent at this tier with bounded exposure before committing fully?
This isn't about how often the agent has been right. It's about what evidence that record actually covers — and what changes at the new level of authority.
SO, WHAT WOULD I DO?
None of these is free — and none is obviously the responsible one.
A acts on a real record and solves a real backlog cost — but treats a track record built entirely below $50K as if it settles a question about $150K. B avoids authorizing an untested tier — but keeps the backlog and its competitive-risk cost running indefinitely. C looks like a reasonable hedge — but customer tenure was never validated as a risk proxy for deal size. D attacks the delay without expanding authority — but assumes the six-day queue is mostly idle time, which may not be true.
A performance record can inform this decision. It can't make it for you.
THE SALES JUDGMENT TAKEAWAY
Reliable enough to trust more isn't automatically authorized further.
Ask what the evidence actually covers.
Ask what being wrong costs at the new level, not just how often it happens.
A performance record can inform that.
It can't decide it for you.
BRING THE CHALLENGE TO YOUR TEAM
Make judgment visible.
Judgment Challenges can be explored with your team through a private 90-minute Judgment Lab, turning the scenario into a practical conversation about AI autonomy, exposure, and how much authority a track record actually earns.