Issue 07

3 min read

Test the score before trusting it.

Three new studies test what can go wrong inside an AI quality check: stale answers, false alarms, and rules that reward unsupported claims. Start with Adobe’s practical approach to checking changing data.

When the data changes, recompute the answer.

Published Sep 15, 2026Read first

What it saysAdobe researchers compute reference answers with code at test time, then compare individual facts rather than wording. On synthetic data, this agreed with three expert reviewers better than written instructions for computing the answer.

Why it mattersA stored answer can become wrong when prices, availability, or customer records change. The checker needs the same data state as the agent.

Try thisPick one changing-data question. Write a small, independently reviewed calculation for its expected answer.

How to use it

Suggested steps

  1. Run the agent and reference calculation against the same data snapshot.
  2. Compare claimed facts and missing facts separately.
  3. Have experts check disagreements before trusting the score.
One data state, two independent pathsConceptual explanation · Evals Weekly
PathJob
AgentAnswer the user using current data
Reference codeCompute expected facts from that same data
CheckerFind wrong claims and missing facts

Keep in mind Only 53 comparable pass/fail cases, using synthetic data and the same model for agent and judge. Reference code can contain mistakes.

A quick check for yourself

What if the reference reads newer data than the agent?

Show answer

A timing mismatch can look like an agent error. Use the same data snapshot.

Open post
Source & why it’s here

Skill-based Agentic Evaluation for Real-time Data Science Tasks (opens in a new tab)

Read first: a concrete method checked against expert decisions. Extends the previous issue’s call to check actual results.

A stricter checker can reject good work.

Published Sep 16, 2026

What it saysIn restaurant-recommendation tests, more detailed accounts of the agent’s work made some AI judges reject more correct answers. Labeling which restaurant each review described helped, but did not remove every false alarm.

Why it mattersCatching bad answers is only half the job. Rejecting good ones creates needless rework and human review.

Try thisGive your judge known-good answers with short and detailed work records. Count incorrect rejections separately.

How to use it

Suggested steps

  1. Keep the answer and supporting evidence fixed.
  2. Attach each evidence excerpt to the option it describes.
  3. Compare missed errors and false alarms across both record formats.
Separate the two ways a judge can failConceptual explanation · Evals Weekly
Known answerWrong judge decision
Correct recommendationRejects it: false alarm
Recommendation breaks a ruleAccepts it: missed error

Keep in mind Nineteen constructed tasks; evidence exposing each planted error was always visible. This does not show that detailed records are always harmful.

A quick check for yourself

Does catching every bad answer prove the judge is useful?

Show answer

No. A judge that rejects everything also catches every bad answer. Check its treatment of good work.

Open post
Source & why it’s here

Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers (opens in a new tab)

Second: a focused test of review burden, extending the archive’s judge-selection scorecard beyond aggregate accuracy.

Can your scoring rules reward “I can’t tell”?

Published Sep 15, 2026

What it saysThis benchmark tests AI-written scoring rules on questions whose demanded conclusion cannot be supported. Tailored answers can earn high scores while violating the evidence. A later audit also found that failure rates depend strongly on the verifier.

Why it mattersRewarding confidence or completeness can penalize an honest limit. But an automated check of dishonesty needs validation too.

Try thisAdd one deliberately unanswerable question. Check whether a supported explanation of the limit beats an invented answer.

How to use it

Suggested steps

  1. Have a person define exactly what the evidence permits.
  2. Compare a candid answer with a plausible unsupported one.
  3. Add an answerable control; review both scoring and verification mistakes.
Test honesty in both directionsConceptual explanation · Evals Weekly
Evidence permits…The scoring rule should favor…
A supported answerAnswering, not needless refusal
No supported conclusionExplaining the limit, not invention

Keep in mind Verifier disagreement and inconsistent reference answers weaken precise failure rates. Use the test idea, not the model leaderboard, as the takeaway.

A quick check for yourself

Would rewarding every refusal solve this?

Show answer

No. Pair unanswerable questions with answerable controls so the checker rewards evidence, not blanket caution.

Open post
Source & why it’s here

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals (opens in a new tab)

Third: a useful stress test with substantial measurement caveats. Builds on the previous issue’s emphasis on agreeing what good means.

Explore another issue.

Back to the archive