Send uncertain scores to a stronger checker.
Published Sep 22, 2026
What it saysA cheap first-pass judge accepted confident decisions and sent uncertain ones to a stronger judge. On pooled public tasks, one threshold escalated 34% of cases and reached 91.3% accuracy versus 91.7% for always using the stronger judge, at 47% of its fee.
Why it mattersA calibrated escalation rule can reduce evaluation cost without treating every case as equally easy. Confidence is useful for routing, not proof that a verdict is right.
Try thisTest a confidence threshold on local human-labeled examples. Track both accuracy and how many cases require escalation.
How to use it
Suggested steps
- Choose the threshold on local examples and freeze it before testing.
- Escalate invalid outputs, uncertain cases, and critical decisions.
- Report error recall, false alarms, and review load for each task group.
| Observed case | Route |
|---|---|
| High confidence on a validated task | Accept first-pass verdict |
| Low confidence or critical decision | Use stronger check or human review |
| Known weak task or style | Bypass the first pass |
Keep in mind The first-pass judge is proprietary. Human adjudication covered 183 disagreement-selected items, used one author as annotator, and left 36 labels undecided. The threshold failed to transfer to every task.
A quick check for yourself
Is high confidence the same as correctness?
Show answer
No. It is a routing signal that must be checked on your own tasks.
Source & why it’s here
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure (opens in a new tab)
Third: a useful cost-and-review pattern with important evidence limits.