A stricter checker can reject good work.
Published Sep 16, 2026
What it saysIn restaurant-recommendation tests, more detailed accounts of the agent’s work made some AI judges reject more correct answers. Labeling which restaurant each review described helped, but did not remove every false alarm.
Why it mattersCatching bad answers is only half the job. Rejecting good ones creates needless rework and human review.
Try thisGive your judge known-good answers with short and detailed work records. Count incorrect rejections separately.
How to use it
Suggested steps
- Keep the answer and supporting evidence fixed.
- Attach each evidence excerpt to the option it describes.
- Compare missed errors and false alarms across both record formats.
| Known answer | Wrong judge decision |
|---|---|
| Correct recommendation | Rejects it: false alarm |
| Recommendation breaks a rule | Accepts it: missed error |
Keep in mind Nineteen constructed tasks; evidence exposing each planted error was always visible. This does not show that detailed records are always harmful.
A quick check for yourself
Does catching every bad answer prove the judge is useful?
Show answer
No. A judge that rejects everything also catches every bad answer. Check its treatment of good work.
Source & why it’s here
Second: a focused test of review burden, extending the archive’s judge-selection scorecard beyond aggregate accuracy.