SALES JUDGMENT CHALLENGE 008

THE
BOUNDARY

An AI deal-approval agent has a 98% clean-outcome rate on renewals up to $50K. Leadership wants to triple its authority to $150K — but every one of those approvals was a smaller deal, and a mistake at the new tier costs three times more.

What would you do?

THE SITUATION

The track record is real. It was never measured at this tier.

An AI deal-approval agent is authorized to independently approve standard renewal contracts up to $50K in annual value — no human sign-off required. Over 18 months and 2,200 approvals, 98% completed cleanly: no disputes, no clawbacks, no compliance flags.

Renewals between $50K and $150K still require human review, and that queue now averages six business days. Of the deals sitting in that queue, 12% have seen the customer request an extension, delay signing, or start engaging a competitor during the wait.

Every one of the 2,200 approvals behind the 98% figure was a deal under $50K. Deals above $50K also tend to carry more negotiated terms, though nobody has established whether that changes how the agent decides.

Even if the agent stayed exactly as reliable at $150K, a mistake there costs three times more.

MAKE THE CALL

What do you do?

Choose before you continue.

A

Raise the ceiling to $150K.

The clean-outcome rate is real and the backlog cost is real — extend the authority the AI has earned.

B

Keep the ceiling at $50K.

Neither the track record nor the higher stakes at this tier have been tested — don't extend authority into either unknown.

C

Raise it for loyal customers.

Combine the size threshold with a relationship signal as a proxy for lower risk.

D

Compress the review SLA.

Cut the six-day queue to one day without expanding AI authority.

WATCH THE CHALLENGE

Coming December 14.

The video for this challenge publishes December 14. Check back then — or work through the situation and questions below in the meantime.

THE JUDGMENT PROBLEM

Reliability and authority are not the same question.

This isn't "has the AI earned more trust?" On the numbers, clearly yes.

The harder question a strong track record can't answer by itself is: even if the AI stayed exactly as reliable at a higher tier, does more reliability automatically justify more authority — when what's at stake per mistake has also changed?

A 98% clean-outcome rate says something about how often the agent gets it right. It says nothing about what should happen with that information once the cost of the 2% changes.

A system can become more reliable without becoming more appropriate to authorize at every level of consequence. Those are two separate judgments.

THE AUTHORITY PROBLEM

A performance record can inform the decision. It can't make it for you.

Even a perfectly transferred 98% wouldn't settle whether $150K of unsupervised exposure is a boundary anyone should accept.

AI can hit 98% consistently. Judgment determines whether 98% is enough at this level of exposure.

THE JUDGMENT DIFFERENCE

Reliable is not the same as authorized.

THE RECORD SHOWS

98% clean, under $50K.

JUDGMENT ASKS

What does being wrong cost at $150K?

The rate might transfer perfectly. The stakes already have.

BETTER QUESTIONS

Before you raise the ceiling, ask:

01

Would you accept this same rate at $150K if it had been measured there directly?

02

Do the negotiated terms common above $50K actually change what the agent is deciding?

03

Is there a way to test the agent at this tier with bounded exposure before committing fully?

This isn't about how often the agent has been right. It's about what evidence that record actually covers — and what changes at the new level of authority.

SO, WHAT WOULD I DO?

It depends.

None of these is free — and none is obviously the responsible one.

A acts on a real record and solves a real backlog cost — but treats a track record built entirely below $50K as if it settles a question about $150K. B avoids authorizing an untested tier — but keeps the backlog and its competitive-risk cost running indefinitely. C looks like a reasonable hedge — but customer tenure was never validated as a risk proxy for deal size. D attacks the delay without expanding authority — but assumes the six-day queue is mostly idle time, which may not be true.

A performance record can inform this decision. It can't make it for you.

THE SALES JUDGMENT TAKEAWAY

01

Reliable enough to trust more isn't automatically authorized further.

02

Ask what the evidence actually covers.

03

Ask what being wrong costs at the new level, not just how often it happens.

A performance record can inform that.
It can't decide it for you.

BRING THE CHALLENGE TO YOUR TEAM

Make judgment visible.

Judgment Challenges can be explored with your team through a private 90-minute Judgment Lab, turning the scenario into a practical conversation about AI autonomy, exposure, and how much authority a track record actually earns.