Issue 04

3 min read

Does your checker notice the right things?

A consistent score can still miss a real mistake. These papers explain how to test what a checker notices—and what it wrongly rewards.

A steady score can still miss a mistake.

Published August 25, 2026Read first

What it saysThe researchers test two abilities: ignoring edits that do not change quality, and noticing edits that do. Judges could look stable while missing meaningful differences.

Why it mattersConsistency alone does not prove that a checker catches the problem you care about.

Try thisMake two copies of an answer: rephrase one and add a factual mistake to the other.

How to use it

Suggested steps

  1. Make two copies of an answer: rephrase one and add a factual mistake to the other.
  2. Check that rephrasing leaves its accuracy score unchanged.
  3. Check that the introduced mistake lowers the score.
Two changes, two expected responsesConceptual explanation · Evals Weekly
CheckWhat to look for
Only the wording changesAccuracy score stays the same
A fact becomes wrongAccuracy score gets worse

Keep in mind The result depends on the kinds of edits tested and the chosen scoring rule.

A quick check for yourself

What is the practical lesson?

Show answer

Consistency alone does not prove that a checker catches the problem you care about.

Open post
Source & why it’s here

A Judge Should Know What Changed (opens in a new tab)

Original issue reading order, preserved from the archive.

Approve each kind of check separately.

Published August 25, 2026

What it saysAI judges were compared with people reviewing voice conversations. Reliability depended on the industry, the question being scored, and the setup; some checks were unsuitable for full automation.

Why it mattersA checker that handles one question well may fail at another.

Try thisCompare human and AI reviews for each quality question separately.

How to use it

Suggested steps

  1. Compare human and AI reviews for each quality question separately.
  2. List the important mistakes the AI misses.
  3. Keep unreliable checks with people; sample the rest regularly.
Choose the review pathConceptual explanation · Evals Weekly
CheckWhat to look for
Reliable check, low stakesAutomate with regular sampling
Important mistakes still missedKeep human review

Keep in mind Findings from the tested conversations and models may not transfer to your calls.

A quick check for yourself

What is the practical lesson?

Show answer

A checker that handles one question well may fail at another.

Open post
Source & why it’s here

Benchmarking LLM Judges for Voice-Agent Evaluation (opens in a new tab)

Original issue reading order, preserved from the archive.

Good writing should not earn a better accuracy score.

Published August 24, 2026

What it saysA score for one quality can be influenced by another. This paper measures that problem and tests removing irrelevant evidence from the checker’s explanation.

Why it mattersSeparate score names do not guarantee separate judgments.

Try thisRephrase an answer without changing any facts.

How to use it

Suggested steps

  1. Rephrase an answer without changing any facts.
  2. Check whether its accuracy score moves.
  3. If it does, tighten the rule and test other unrelated changes.
Keep the question specificConceptual explanation · Evals Weekly
CheckWhat to look for
Writing becomes clearerClarity score may improve
Facts are unchangedAccuracy score should hold

Keep in mind Some qualities genuinely interact. Investigate whether each score change makes sense.

A quick check for yourself

What is the practical lesson?

Show answer

Separate score names do not guarantee separate judgments.

Open post
Source & why it’s here

Inter-dimension Dependence (opens in a new tab)

Original issue reading order, preserved from the archive.

The AI may already have found the right answer—and discarded it.

Published August 26, 2026

What it saysThe study separates producing possible answers, recognizing a correct one, and choosing the final response. Changing the selection rule improved results without generating new answers.

Why it mattersFind the failing step before changing the model.

Try thisSave the alternatives your assistant considered for one failed task.

How to use it

Suggested steps

  1. Save the alternatives your assistant considered for one failed task.
  2. Check whether a correct answer was present and recognized.
  3. Fix generation, checking, or selection based on where it was lost.
Where did the right answer disappear?Conceptual explanation · Evals Weekly
CheckWhat to look for
No correct option was generatedImprove the available answers
Correct option was discardedImprove checking or selection

Keep in mind The reported improvements depend on the answer pools and tasks in the study.

A quick check for yourself

What is the practical lesson?

Show answer

Find the failing step before changing the model.

Open post
Source & why it’s here

Candidate Supply and Answer Selection (opens in a new tab)

Original issue reading order, preserved from the archive.

Explore another issue.

Back to the archive