From Screen broadly. Confirm with real outcomes.Issue published Sep 25, 2026

Ten judges may share one blind spot.

Published Sep 18, 2026

What it saysIn the main ten-judge test, shared errors reduced the independent information to roughly 3.5 judges. In up to 28% of system comparisons, a claimed significant difference disappeared after the analysis accounted for those shared mistakes.

Why it mattersAdding more judges can create false confidence when they fail on the same examples. Using models from different providers does not guarantee independent evidence.

Try thisRun every judge on about 100 trusted examples. Look for cases that fool several judges before choosing a voting rule.

How to use it

Suggested steps

  1. Use the same trusted set and hidden labels for every judge.
  2. Group shared mistakes by failure type, not only by model.
  3. Choose the voting rule on one set; verify it on held-out examples.
Conceptual view of shared errorsConceptual explanation · Evals Weekly
What you seeWhat may be true
8 of 10 judges agreeSeveral copied the same blind spot
Trusted examples expose clustersAgreement can be weighted more honestly

Keep in mind The main bank used five small open models with two prompt styles. Controlled tests focused on position bias, and the numerical results should not be assumed to transfer to another judge bank.

A quick check for yourself

Does using different model providers make judge votes independent?

Show answer

Not necessarily. Check shared mistakes on trusted examples.

Source & why it’s here