From Test the score before trusting it.Issue published Sep 18, 2026

A stricter checker can reject good work.

Published Sep 16, 2026

What it saysIn restaurant-recommendation tests, more detailed accounts of the agent’s work made some AI judges reject more correct answers. Labeling which restaurant each review described helped, but did not remove every false alarm.

Why it mattersCatching bad answers is only half the job. Rejecting good ones creates needless rework and human review.

Try thisGive your judge known-good answers with short and detailed work records. Count incorrect rejections separately.

How to use it

Suggested steps

  1. Keep the answer and supporting evidence fixed.
  2. Attach each evidence excerpt to the option it describes.
  3. Compare missed errors and false alarms across both record formats.
Separate the two ways a judge can failConceptual explanation · Evals Weekly
Known answerWrong judge decision
Correct recommendationRejects it: false alarm
Recommendation breaks a ruleAccepts it: missed error

Keep in mind Nineteen constructed tasks; evidence exposing each planted error was always visible. This does not show that detailed records are always harmful.

A quick check for yourself

Does catching every bad answer prove the judge is useful?

Show answer

No. A judge that rejects everything also catches every bad answer. Check its treatment of good work.

Source & why it’s here

Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers (opens in a new tab)

Second: a focused test of review burden, extending the archive’s judge-selection scorecard beyond aggregate accuracy.