From Screen broadly. Confirm with real outcomes.Issue published Sep 25, 2026

Send uncertain scores to a stronger checker.

Published Sep 22, 2026

What it saysA cheap first-pass judge accepted confident decisions and sent uncertain ones to a stronger judge. On pooled public tasks, one threshold escalated 34% of cases and reached 91.3% accuracy versus 91.7% for always using the stronger judge, at 47% of its fee.

Why it mattersA calibrated escalation rule can reduce evaluation cost without treating every case as equally easy. Confidence is useful for routing, not proof that a verdict is right.

Try thisTest a confidence threshold on local human-labeled examples. Track both accuracy and how many cases require escalation.

How to use it

Suggested steps

  1. Choose the threshold on local examples and freeze it before testing.
  2. Escalate invalid outputs, uncertain cases, and critical decisions.
  3. Report error recall, false alarms, and review load for each task group.
Conceptual escalation ruleConceptual explanation · Evals Weekly
Observed caseRoute
High confidence on a validated taskAccept first-pass verdict
Low confidence or critical decisionUse stronger check or human review
Known weak task or styleBypass the first pass

Keep in mind The first-pass judge is proprietary. Human adjudication covered 183 disagreement-selected items, used one author as annotator, and left 36 labels undecided. The threshold failed to transfer to every task.

A quick check for yourself

Is high confidence the same as correctness?

Show answer

No. It is a routing signal that must be checked on your own tasks.

Source & why it’s here