HR JUDGMENT CHALLENGE 003
THE AI-ADJUSTED
PERFORMANCE REVIEW
Two claims processors, same role. After a complexity adjustment built to fix exactly this kind of complaint, one employee's score is still 15% higher — and nobody's sure whether that's about performance, or about tooling.
What would you do?
THE SITUATION
The score is accurate. The question is what it's evidence of.
Two claims processors work the same role, each paired with an AI assistant. Employee A handles standardized claims; Employee B handles escalated, non-standard ones.
Six months ago, HR replaced the old raw-output metric with a complexity-adjusted productivity score, weighting each claim type by how hard it's historically been to resolve — independent of who or what does the work. It was built to stop penalizing employees for harder caseloads, and it does that.
What the weights don't measure is how much of the work is now done by the AI versus the employee. On A's standardized claims, her AI assistant resolves about 70% autonomously. On B's escalated claims, AI contributes maybe 20% — the rest is her own judgment.
This quarter, A completed 42 cases, B completed 31. After the complexity weighting, A's score is still 15% higher. A promotion decision is due this cycle.
The weights are accurate. Nobody yet knows what the 15% means.
MAKE THE CALL
What do you do?
Choose before you continue.
WATCH THE CHALLENGE
Coming December 7.
The video for this challenge publishes December 7. Check back then — or work through the situation and questions below in the meantime.
THE JUDGMENT PROBLEM
Two questions, one number.
The complexity weighting isn't broken. It measures exactly what it was built to measure — how hard a claim type is — and it does that accurately regardless of who resolves the case.
The question it was never designed to answer is the one HR is now asking it: what does individual performance mean when the individual's output is increasingly produced by a human-AI system, not the individual alone?
A's 15% advantage isn't necessarily a distortion. If the question is who produces more value in the role as currently tooled, the score may be exactly right.
But if the question is who has more individual capability or promotion readiness, the same 15% becomes much weaker evidence. Same number. Two different questions.
THE EVIDENCE PROBLEM
The score can't tell you which question it's answering.
Same number. Two different questions. Only you can decide which one you're asking.
AI can measure productivity perfectly. Judgment determines whether productivity is what you're actually rewarding.
THE JUDGMENT DIFFERENCE
Accurate is not the same as relevant.
Output in the role as tooled.
Is that the same as promotion-ready?
Same 15%. Different question, different answer.
BETTER QUESTIONS
Before you decide the promotion, ask:
Is this about who produces more as currently tooled — or who has the capability for more responsibility?
Would the gap survive if both worked the same mix of cases, with matched AI support?
What does each mistake cost: promoting on tooling, or withholding promotion from someone whose caseload never gave her the chance?
This isn't about whether the metric is accurate. It's about what it's actually evidence of — and whether that's the same thing you're using it to decide.
SO, WHAT WOULD I DO?
Each option is defensible — depending on what the promotion is actually asking.
A treats the score as legitimate evidence of role performance — which it may well be, if that's genuinely the question. B discards real information HR spent months building. C sounds rigorous, but a matched-case trial still has to answer the same underlying question afterward. D separates what the score is good evidence for from what it isn't, but the promotion decision doesn't resolve this cycle.
This isn't about whether the metric is accurate. It's about what it's actually evidence of — and whether that's the same thing you're using it to decide.
THE HR JUDGMENT TAKEAWAY
An improved metric can still answer a different question than you think.
Don't only ask if it's accurate. Ask what it's evidence of.
Know which question the promotion is actually asking.
Don't only ask whether the metric is accurate.
Ask what it's actually evidence of.
BRING THE CHALLENGE TO YOUR TEAM
Make judgment visible.
Judgment Challenges can be explored with your team through a private 90-minute Judgment Lab, turning the scenario into a practical conversation about performance metrics, AI-assisted work, and what a number is actually evidence of.