Can your scoring rules reward “I can’t tell”?
Published Sep 15, 2026
What it saysThis benchmark tests AI-written scoring rules on questions whose demanded conclusion cannot be supported. Tailored answers can earn high scores while violating the evidence. A later audit also found that failure rates depend strongly on the verifier.
Why it mattersRewarding confidence or completeness can penalize an honest limit. But an automated check of dishonesty needs validation too.
Try thisAdd one deliberately unanswerable question. Check whether a supported explanation of the limit beats an invented answer.
How to use it
Suggested steps
- Have a person define exactly what the evidence permits.
- Compare a candid answer with a plausible unsupported one.
- Add an answerable control; review both scoring and verification mistakes.
| Evidence permits… | The scoring rule should favor… |
|---|---|
| A supported answer | Answering, not needless refusal |
| No supported conclusion | Explaining the limit, not invention |
Keep in mind Verifier disagreement and inconsistent reference answers weaken precise failure rates. Use the test idea, not the model leaderboard, as the takeaway.
A quick check for yourself
Would rewarding every refusal solve this?
Show answer
No. Pair unanswerable questions with answerable controls so the checker rewards evidence, not blanket caution.
Source & why it’s here
ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals (opens in a new tab)
Third: a useful stress test with substantial measurement caveats. Builds on the previous issue’s emphasis on agreeing what good means.