Ten judges may share one blind spot.
Published Sep 18, 2026
What it saysIn the main ten-judge test, shared errors reduced the independent information to roughly 3.5 judges. In up to 28% of system comparisons, a claimed significant difference disappeared after the analysis accounted for those shared mistakes.
Why it mattersAdding more judges can create false confidence when they fail on the same examples. Using models from different providers does not guarantee independent evidence.
Try thisRun every judge on about 100 trusted examples. Look for cases that fool several judges before choosing a voting rule.
How to use it
Suggested steps
- Use the same trusted set and hidden labels for every judge.
- Group shared mistakes by failure type, not only by model.
- Choose the voting rule on one set; verify it on held-out examples.
| What you see | What may be true |
|---|---|
| 8 of 10 judges agree | Several copied the same blind spot |
| Trusted examples expose clusters | Agreement can be weighted more honestly |
Keep in mind The main bank used five small open models with two prompt styles. Controlled tests focused on position bias, and the numerical results should not be assumed to transfer to another judge bank.
A quick check for yourself
Does using different model providers make judge votes independent?
Show answer
Not necessarily. Check shared mistakes on trusted examples.
Source & why it’s here
Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus (opens in a new tab)
Second: a direct warning against treating judge votes as independent evidence.