Issue 08September 18–25

3 min read

Screen broadly. Confirm with real outcomes.

Three new studies improve the path from an offline score to a trustworthy release decision. Start with Nubank’s production-tested approach: simulate customer journeys, then let real customer outcomes decide what ships.

Open issue

Test with synthetic customers before real ones.

Published Sep 24, 2026Read first

What it saysNubank compared agent versions through multi-turn synthetic customers and mocked tool responses. A configuration selected after more than 16,000 simulated conversations later improved self-service by 8.82 percentage points in a live test, with no significant change in customer satisfaction.

Why it mattersThis adds a useful layer between hand-written tests and exposing real customers. Simulation finds weak candidates cheaply; customer behavior still confirms the release.

Try thisChoose one failure hypothesis. Run the incumbent and candidate through the same 100 simulated journeys before a small live test.

How to use it

Suggested steps

  1. Keep user scenarios and mocked tool states identical for both versions.
  2. Calibrate each new checker with domain experts before using it to screen.
  3. Only advance clear wins; confirm them with a real product outcome.
Conceptual release ladderConceptual explanation · Evals Weekly
LayerDecision it supports
SimulationIs this candidate safe enough to test?
Expert calibrationDoes the checker apply the intended rule?
Live outcomeDid the customer experience improve?

Keep in mind The synthetic customers wrote much longer messages than real users, and mocked tools did not test backend behavior, latency, saved state, or side effects. Treat simulation as a screen, not proof.

A quick check for yourself

Can a simulated win approve a release by itself?

Show answer

No. It can screen candidates; a real product outcome should confirm the winner.

Open post
Source & why it’s here

Ten judges may share one blind spot.

Published Sep 18, 2026

What it saysIn the main ten-judge test, shared errors reduced the independent information to roughly 3.5 judges. In up to 28% of system comparisons, a claimed significant difference disappeared after the analysis accounted for those shared mistakes.

Why it mattersAdding more judges can create false confidence when they fail on the same examples. Using models from different providers does not guarantee independent evidence.

Try thisRun every judge on about 100 trusted examples. Look for cases that fool several judges before choosing a voting rule.

How to use it

Suggested steps

  1. Use the same trusted set and hidden labels for every judge.
  2. Group shared mistakes by failure type, not only by model.
  3. Choose the voting rule on one set; verify it on held-out examples.
Conceptual view of shared errorsConceptual explanation · Evals Weekly
What you seeWhat may be true
8 of 10 judges agreeSeveral copied the same blind spot
Trusted examples expose clustersAgreement can be weighted more honestly

Keep in mind The main bank used five small open models with two prompt styles. Controlled tests focused on position bias, and the numerical results should not be assumed to transfer to another judge bank.

A quick check for yourself

Does using different model providers make judge votes independent?

Show answer

Not necessarily. Check shared mistakes on trusted examples.

Open post
Source & why it’s here

Send uncertain scores to a stronger checker.

Published Sep 22, 2026

What it saysA cheap first-pass judge accepted confident decisions and sent uncertain ones to a stronger judge. On pooled public tasks, one threshold escalated 34% of cases and reached 91.3% accuracy versus 91.7% for always using the stronger judge, at 47% of its fee.

Why it mattersA calibrated escalation rule can reduce evaluation cost without treating every case as equally easy. Confidence is useful for routing, not proof that a verdict is right.

Try thisTest a confidence threshold on local human-labeled examples. Track both accuracy and how many cases require escalation.

How to use it

Suggested steps

  1. Choose the threshold on local examples and freeze it before testing.
  2. Escalate invalid outputs, uncertain cases, and critical decisions.
  3. Report error recall, false alarms, and review load for each task group.
Conceptual escalation ruleConceptual explanation · Evals Weekly
Observed caseRoute
High confidence on a validated taskAccept first-pass verdict
Low confidence or critical decisionUse stronger check or human review
Known weak task or styleBypass the first pass

Keep in mind The first-pass judge is proprietary. Human adjudication covered 183 disagreement-selected items, used one author as annotator, and left 36 labels undecided. The threshold failed to transfer to every task.

A quick check for yourself

Is high confidence the same as correctness?

Show answer

No. It is a routing signal that must be checked on your own tasks.

Open post
Source & why it’s here
One idea from the previous issue

Why recompute a reference answer when data changes?

Show answer

A stored answer can become stale. The agent and checker need the same data state.

Revisit the issue

You’re caught up.

Browse earlier issues

Keep exploring

Earlier issues

7 issues · 25 reads

07Sep 18, 2026 · 3 reads · 3 min briefTest the score before trusting it.
Open issue

Three new studies test what can go wrong inside an AI quality check: stale answers, false alarms, and rules that reward unsupported claims. Start with Adobe’s practical approach to checking changing data.

When the data changes, recompute the answer.

Published Sep 15, 2026

What it saysAdobe researchers compute reference answers with code at test time, then compare individual facts rather than wording. On synthetic data, this agreed with three expert reviewers better than written instructions for computing the answer.

Why it mattersA stored answer can become wrong when prices, availability, or customer records change. The checker needs the same data state as the agent.

Try thisPick one changing-data question. Write a small, independently reviewed calculation for its expected answer.

How to use it

Suggested steps

  1. Run the agent and reference calculation against the same data snapshot.
  2. Compare claimed facts and missing facts separately.
  3. Have experts check disagreements before trusting the score.
One data state, two independent pathsConceptual explanation · Evals Weekly
PathJob
AgentAnswer the user using current data
Reference codeCompute expected facts from that same data
CheckerFind wrong claims and missing facts

Keep in mind Only 53 comparable pass/fail cases, using synthetic data and the same model for agent and judge. Reference code can contain mistakes.

Open post
Source & why it’s here

Skill-based Agentic Evaluation for Real-time Data Science Tasks (opens in a new tab)

Read first: a concrete method checked against expert decisions. Extends the previous issue’s call to check actual results.

A stricter checker can reject good work.

Published Sep 16, 2026

What it saysIn restaurant-recommendation tests, more detailed accounts of the agent’s work made some AI judges reject more correct answers. Labeling which restaurant each review described helped, but did not remove every false alarm.

Why it mattersCatching bad answers is only half the job. Rejecting good ones creates needless rework and human review.

Try thisGive your judge known-good answers with short and detailed work records. Count incorrect rejections separately.

How to use it

Suggested steps

  1. Keep the answer and supporting evidence fixed.
  2. Attach each evidence excerpt to the option it describes.
  3. Compare missed errors and false alarms across both record formats.
Separate the two ways a judge can failConceptual explanation · Evals Weekly
Known answerWrong judge decision
Correct recommendationRejects it: false alarm
Recommendation breaks a ruleAccepts it: missed error

Keep in mind Nineteen constructed tasks; evidence exposing each planted error was always visible. This does not show that detailed records are always harmful.

Open post
Source & why it’s here

Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers (opens in a new tab)

Second: a focused test of review burden, extending the archive’s judge-selection scorecard beyond aggregate accuracy.

Can your scoring rules reward “I can’t tell”?

Published Sep 15, 2026

What it saysThis benchmark tests AI-written scoring rules on questions whose demanded conclusion cannot be supported. Tailored answers can earn high scores while violating the evidence. A later audit also found that failure rates depend strongly on the verifier.

Why it mattersRewarding confidence or completeness can penalize an honest limit. But an automated check of dishonesty needs validation too.

Try thisAdd one deliberately unanswerable question. Check whether a supported explanation of the limit beats an invented answer.

How to use it

Suggested steps

  1. Have a person define exactly what the evidence permits.
  2. Compare a candid answer with a plausible unsupported one.
  3. Add an answerable control; review both scoring and verification mistakes.
Test honesty in both directionsConceptual explanation · Evals Weekly
Evidence permits…The scoring rule should favor…
A supported answerAnswering, not needless refusal
No supported conclusionExplaining the limit, not invention

Keep in mind Verifier disagreement and inconsistent reference answers weaken precise failure rates. Use the test idea, not the model leaderboard, as the takeaway.

Open post
Source & why it’s here

ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals (opens in a new tab)

Third: a useful stress test with substantial measurement caveats. Builds on the previous issue’s emphasis on agreeing what good means.

06Sep 13, 2026 · 3 reads · 3 min briefWho checks the AI checker?
Open issue

Netflix’s latest post shows why AI checks need regular human review. Two earlier reads explain how to agree on the rules and test what actually happened.

An AI checker needs checkups, too.

Published Sep 4, 2026

What it saysNetflix uses AI to check explanations of show recommendations. Experts supply examples and reasons; the team adjusts the checker, then keeps comparing its decisions with fresh human reviews.

Why it mattersA checker that worked at launch can become unreliable as the content changes. Matching a human’s verdict is not enough if the reason is wrong.

Try thisReview a small sample of recent answers. Record where people and the checker disagree, including why.

How to use it

Suggested steps

  1. Write one clear pass/fail rule with examples.
  2. Compare human decisions and reasons with the checker’s.
  3. Fix disagreements; retest on examples kept aside. Repeat regularly.
Keep the check connected to peopleConceptual explanation · Evals Weekly
WhenWhat to check
Before launchDoes it agree with experts, for the right reason?
After launchDoes that still hold on fresh examples?

Keep in mind A Netflix case study; results may differ for other products.

Open post
Source & why it’s here

Agree on “good” before asking AI to score it.

Published Apr 29, 2026

What it saysPeople can use the same scoring rule and mean different things. MultEval lets a team compare disagreements, revise rules together, and keep the examples behind each decision.

Why it mattersAn AI checker cannot resolve a quality standard your team has never agreed on.

Try thisAsk two colleagues to review the same answers separately. Start the discussion with their disagreements.

How to use it

Suggested steps

  1. Choose examples that split opinion.
  2. Ask what each person understood the rule to mean.
  3. Rewrite the rule together; save a clear example beside it.
Find the ambiguityConceptual explanation · Evals Weekly
Vague ruleA clearer question
“Be helpful”Did the answer resolve the user’s request?
“Be accurate”Which claim does the source support?

Keep in mind A study of collaboration and a prototype, not proof of a universal accuracy gain.

Open post
Source & why it’s here

Check the result, not just the reply.

Published Jan 9, 2026

What it saysAn AI assistant can say it completed a task without completing it. Anthropic recommends combining software checks, AI judgments, and human review, chosen for the task.

Why it mattersA convincing “done” message tells you little about whether a booking, file edit, or support request succeeded.

Try thisFor one task, define the evidence that would prove it worked. Check that evidence directly.

How to use it

Suggested steps

  1. Turn a real failure into a repeatable test.
  2. Check the saved result with code where possible.
  3. Use people to check subjective quality; repeat the test after changes.
Match the check to the questionConceptual explanation · Evals Weekly
QuestionUseful check
Was the change saved?Inspect the saved record
Was the answer useful?Compare with human review

Keep in mind Practice guidance; each product still needs its own success criteria.

Open post
Source & why it’s here

Demystifying evals for AI agents (opens in a new tab)

A practical starting point for testing an entire task.

05Sep 4, 2026 · 3 reads · 3 min briefReuse the work. Verify the result.
Open issue

Three ways to make AI testing more useful: reuse earlier reviews, check changing facts, and verify exactly what an assistant changed.

Stop paying to check the same evidence twice.

Published September 2, 2026

What it saysWhen comparing systems that find documents, many results overlap. This method saves earlier judgments and reviews only new documents, then compares every system using the updated shared collection.

Why it mattersEach round of testing can make the next round cheaper.

Try thisCompare the documents returned by two search versions.

How to use it

Suggested steps

  1. Compare the documents returned by two search versions.
  2. Keep earlier judgments when the document and scoring rule are unchanged.
  3. Judge the new documents; compare both versions on the shared collection.
What needs a fresh review?Conceptual explanation · Evals Weekly
CheckWhat to look for
Same document, same ruleReuse the judgment
New document or changed ruleReview again

Keep in mind Recheck judgments when the content or quality rule changes. Report uncertainty around close results.

Open post
Source & why it’s here

Incremental Pooled LLM Evaluation (opens in a new tab)

Original issue reading order, preserved from the archive.

When the facts change, does the AI change its plan?

Published September 3, 2026

What it saysThis test gives assistants conflicting information from users, stored knowledge, and tools. The tested models did not reliably handle every kind of conflict.

Why it mattersMentioning a new fact is not enough. The assistant must use it in its next action.

Try thisCreate a test where an old availability result conflicts with a fresh one.

How to use it

Suggested steps

  1. Create a test where an old availability result conflicts with a fresh one.
  2. Check which result the assistant uses when it acts.
  3. Inspect the resulting record and review any mistaken decisions.
A changed fact should change the actionConceptual explanation · Evals Weekly
CheckWhat to look for
Old result: availableAssistant has an initial plan
Fresh result: unavailablePlan and action must update

Keep in mind A controlled test does not cover every conflict a real product will encounter.

Open post
Source & why it’s here

KC-Bench (opens in a new tab)

Original issue reading order, preserved from the archive.

Let AI propose the change. Make code apply it precisely.

Published August 31, 2026

What it saysIn a study of editing server configuration files, apparently successful edits could change the wrong thing. The proposed approach separates deciding what to change from applying the exact edit.

Why it mattersA correct request can still fail during execution.

Try thisChoose one action where the assistant changes a saved record.

How to use it

Suggested steps

  1. Choose one action where the assistant changes a saved record.
  2. Have it specify the record, field, and new value.
  3. Use code to validate and apply that change; stop if the target is unclear.
Two jobs, two checksConceptual explanation · Evals Weekly
CheckWhat to look for
Decide what should changeReview the proposed action
Apply the changeCheck the exact saved result

Keep in mind The study concerns a specific file-editing task. The proposed change still needs to be appropriate.

Open post
Source & why it’s here

Don’t Let the Model Write the YAML (opens in a new tab)

Original issue reading order, preserved from the archive.

04Aug 28, 2026 · 4 reads · 3 min briefDoes your checker notice the right things?
Open issue

A consistent score can still miss a real mistake. These papers explain how to test what a checker notices—and what it wrongly rewards.

A steady score can still miss a mistake.

Published August 25, 2026

What it saysThe researchers test two abilities: ignoring edits that do not change quality, and noticing edits that do. Judges could look stable while missing meaningful differences.

Why it mattersConsistency alone does not prove that a checker catches the problem you care about.

Try thisMake two copies of an answer: rephrase one and add a factual mistake to the other.

How to use it

Suggested steps

  1. Make two copies of an answer: rephrase one and add a factual mistake to the other.
  2. Check that rephrasing leaves its accuracy score unchanged.
  3. Check that the introduced mistake lowers the score.
Two changes, two expected responsesConceptual explanation · Evals Weekly
CheckWhat to look for
Only the wording changesAccuracy score stays the same
A fact becomes wrongAccuracy score gets worse

Keep in mind The result depends on the kinds of edits tested and the chosen scoring rule.

Open post
Source & why it’s here

A Judge Should Know What Changed (opens in a new tab)

Original issue reading order, preserved from the archive.

Approve each kind of check separately.

Published August 25, 2026

What it saysAI judges were compared with people reviewing voice conversations. Reliability depended on the industry, the question being scored, and the setup; some checks were unsuitable for full automation.

Why it mattersA checker that handles one question well may fail at another.

Try thisCompare human and AI reviews for each quality question separately.

How to use it

Suggested steps

  1. Compare human and AI reviews for each quality question separately.
  2. List the important mistakes the AI misses.
  3. Keep unreliable checks with people; sample the rest regularly.
Choose the review pathConceptual explanation · Evals Weekly
CheckWhat to look for
Reliable check, low stakesAutomate with regular sampling
Important mistakes still missedKeep human review

Keep in mind Findings from the tested conversations and models may not transfer to your calls.

Open post
Source & why it’s here

Benchmarking LLM Judges for Voice-Agent Evaluation (opens in a new tab)

Original issue reading order, preserved from the archive.

Good writing should not earn a better accuracy score.

Published August 24, 2026

What it saysA score for one quality can be influenced by another. This paper measures that problem and tests removing irrelevant evidence from the checker’s explanation.

Why it mattersSeparate score names do not guarantee separate judgments.

Try thisRephrase an answer without changing any facts.

How to use it

Suggested steps

  1. Rephrase an answer without changing any facts.
  2. Check whether its accuracy score moves.
  3. If it does, tighten the rule and test other unrelated changes.
Keep the question specificConceptual explanation · Evals Weekly
CheckWhat to look for
Writing becomes clearerClarity score may improve
Facts are unchangedAccuracy score should hold

Keep in mind Some qualities genuinely interact. Investigate whether each score change makes sense.

Open post
Source & why it’s here

Inter-dimension Dependence (opens in a new tab)

Original issue reading order, preserved from the archive.

The AI may already have found the right answer—and discarded it.

Published August 26, 2026

What it saysThe study separates producing possible answers, recognizing a correct one, and choosing the final response. Changing the selection rule improved results without generating new answers.

Why it mattersFind the failing step before changing the model.

Try thisSave the alternatives your assistant considered for one failed task.

How to use it

Suggested steps

  1. Save the alternatives your assistant considered for one failed task.
  2. Check whether a correct answer was present and recognized.
  3. Fix generation, checking, or selection based on where it was lost.
Where did the right answer disappear?Conceptual explanation · Evals Weekly
CheckWhat to look for
No correct option was generatedImprove the available answers
Correct option was discardedImprove checking or selection

Keep in mind The reported improvements depend on the answer pools and tasks in the study.

Open post
Source & why it’s here

Candidate Supply and Answer Selection (opens in a new tab)

Original issue reading order, preserved from the archive.

03Aug 21, 2026 · 4 reads · 3 min briefThe right score needs the right evidence.
Open issue

File versions, scoring scales, and longer tasks can all change the result. Four papers show what a single overall score leaves out.

Make sure the AI and its checker see the same file.

Published August 18, 2026

What it saysThis work makes file state explicit. What an assistant can read, edit, and submit affects its measured performance; different representations can otherwise refer to different versions.

Why it mattersA version mismatch can look like an AI mistake or hide one.

Try thisRecord which file version the assistant reads and edits.

How to use it

Suggested steps

  1. Record which file version the assistant reads and edits.
  2. Link the readable text and original file to the same version.
  3. Check the submitted file, including the changes actually saved.
Keep the evidence togetherConceptual explanation · Evals Weekly
CheckWhat to look for
What the assistant readsThe correct source version
What the checker reviewsThe actual submitted version

Keep in mind Better file access can improve performance without changing the underlying model.

Open post
Source & why it’s here

StagedWorkspace (opens in a new tab)

Original issue reading order, preserved from the archive.

Getting the order right does not make the scores right.

Published August 18, 2026

What it saysAutomatic scoring ranked the compared systems in the same order as the human jury, but compressed the differences between weak and strong results.

Why it mattersA score can pick the right winner and still set the wrong pass mark.

Try thisCompare AI and human scores for weak, average, and strong answers.

How to use it

Suggested steps

  1. Compare AI and human scores for weak, average, and strong answers.
  2. Check the size of the gaps as well as the order.
  3. Review examples near your release pass mark with people.
Two different questionsConceptual explanation · Evals Weekly
CheckWhat to look for
Who performed better?Compare the ranking
How much better?Compare the scoring scale

Keep in mind The study uses competition submissions. Your product needs its own human comparison.

Open post
Source & why it’s here

The IOL-AI Challenge (opens in a new tab)

Original issue reading order, preserved from the archive.

Finishing a task is only the first check.

Published August 19, 2026

What it saysThis benchmark follows assistants through hundreds of decisions in a long simulation. All tested models finished, yet the quality of their outcomes differed and became clearer late in the task.

Why it mattersA completion count can hide weak decisions along the way.

Try thisChoose one workflow with several linked decisions.

How to use it

Suggested steps

  1. Choose one workflow with several linked decisions.
  2. Check intermediate decisions as well as the final outcome.
  3. Repeat long enough to see whether early mistakes accumulate.
Look beyond completionConceptual explanation · Evals Weekly
CheckWhat to look for
Did it finish?Basic completion
Did its decisions work out?Quality over the whole task

Keep in mind A simulated environment cannot establish how the same behavior will perform in the real world.

Open post
Source & why it’s here

FM-Bench (opens in a new tab)

Original issue reading order, preserved from the archive.

A better total can hide a worse trade-off.

Published August 15, 2026

What it saysThis study compares an assistant with professional estimators on real building plans. It checks how much material was identified separately from how accurate the quantities were.

Why it mattersFinding more items and measuring them correctly are different abilities.

Try thisSplit your overall score into the qualities that matter.

How to use it

Suggested steps

  1. Split your overall score into the qualities that matter.
  2. Compare the assistant with expert-reviewed examples on each quality.
  3. Inspect whether a higher total comes with an unacceptable weakness.
Show both sides of the resultConceptual explanation · Evals Weekly
CheckWhat to look for
CoverageWere the required items found?
AccuracyWere their quantities right?

Keep in mind The findings concern a specific estimation task and expert reference set.

Open post
Source & why it’s here

Handoff-H1 (opens in a new tab)

Original issue reading order, preserved from the archive.

02Aug 14, 2026 · 4 reads · 3 min briefOne score cannot tell the whole story.
Open issue

A good average can hide missed mistakes, wasteful steps, and unreliable behavior. Start with the failures that matter to your users.

Choose a checker by the mistakes it catches.

Published August 9, 2026

What it saysJudges with similar overall accuracy differed sharply in detecting unwanted side effects. Catching more problems also sent more work to human reviewers.

Why it mattersA good average can hide the very failures you need to catch.

Try thisList your most important failure types.

How to use it

Suggested steps

  1. List your most important failure types.
  2. Check how many real problems each judge catches and how many false alarms it creates.
  3. Choose a setting your review team can handle; recheck after changes.
Measure both costsConceptual explanation · Evals Weekly
CheckWhat to look for
Missed mistakesBad answers that passed
False alarmsGood answers sent for review

Keep in mind Higher detection can create more review work. The best trade-off depends on the product.

Open post
Source & why it’s here

Selecting LLM Judges for Agent Evaluation Pipelines (opens in a new tab)

Original issue reading order, preserved from the archive.

A better-looking answer can come from a broken process.

Published August 9, 2026

What it saysIn analytics tasks, a stronger configuration answered more questions using real data while still showing unreliable choices during its work. The final answer alone missed these issues.

Why it mattersCheck how the assistant reached an answer as well as the answer itself.

Try thisSave the data and tool calls used for a task.

How to use it

Suggested steps

  1. Save the data and tool calls used for a task.
  2. Check that the assistant interpreted the right tables and respected its limits.
  3. Repeat the task to see whether the process stays reliable.
Check the answer and the pathConceptual explanation · Evals Weekly
CheckWhat to look for
Final responseIs the answer supported?
Work recordDid it use the right data and steps?

Keep in mind These observations come from an enterprise analytics setup, not every assistant.

Open post
Source & why it’s here

Evaluating Enterprise Analytics Agents (opens in a new tab)

Original issue reading order, preserved from the archive.

The same model can behave very differently in a different setup.

Published August 9, 2026

What it saysChanging the tools and control logic around a coding model changed its cost and failure patterns, even when task success rates were fairly close.

Why it mattersA model score is also a score for the setup around it.

Try thisRecord the model, instructions, tools, retry limits, and stopping rules for each test.

How to use it

Suggested steps

  1. Record the model, instructions, tools, retry limits, and stopping rules for each test.
  2. Change one part at a time.
  3. Compare success, cost, and failure type together.
Make comparisons fairConceptual explanation · Evals Weekly
CheckWhat to look for
Same model, different toolsA different system is being tested
Same model and setupResults are easier to compare

Keep in mind Results from coding tasks may not transfer to other products.

Open post
Source & why it’s here

The Scaffold Effect in Coding Agents (opens in a new tab)

Original issue reading order, preserved from the archive.

Before using a leaderboard, ask what it really tests.

Published August 9, 2026

What it saysModels changed position across different agent benchmarks. A broad score can mix reasoning, tool use, and other abilities in ways that hide what improved.

Why it mattersChoose tests that reflect the work your users need done.

Try thisName the specific ability you need to measure.

How to use it

Suggested steps

  1. Name the specific ability you need to measure.
  2. Check which tasks in a benchmark actually require that ability.
  3. Test critical abilities separately before combining scores.
Start with the intended useConceptual explanation · Evals Weekly
CheckWhat to look for
Your user’s taskThe ability that matters
Benchmark taskEvidence it measures that ability

Keep in mind Different rankings are not automatically an error: different tests may measure different abilities.

Open post
Source & why it’s here

Construct Validity Failures in Agentic AI Benchmarks (opens in a new tab)

Original issue reading order, preserved from the archive.

01Aug 7, 2026 · 4 reads · 3 min briefBetter answers, not prettier words.
Open issue

Checkers can be distracted by writing style. This issue also looks at learning from real work, testing efficiently, and choosing useful examples.

Polished wording can fool an AI checker.

Published August 3 · revised August 4, 2026

What it saysWhen the same substance was written in different styles, judges changed their preferences. A style-aware checking step helped separate the content from its presentation.

Why it mattersAn answer should not get credit for sounding better when its facts are unchanged.

Try thisWrite two versions of an answer with the same facts but different styles.

How to use it

Suggested steps

  1. Write two versions of an answer with the same facts but different styles.
  2. Check whether the factual score changes.
  3. Separate the accuracy question from the writing-quality question.
One answer, two questionsConceptual explanation · Evals Weekly
CheckWhat to look for
Are its facts correct?Check the evidence
Is it easy to read?Check the writing

Keep in mind Style matters for readability; the point is to stop it distorting unrelated scores.

Open post
Source & why it’s here

Style Wins, Substance Loses (opens in a new tab)

Original issue reading order, preserved from the archive.

Write quality rules from real examples of good work.

Published August 4, 2026

What it saysThis study builds scoring rules from professional deliverables. Those rules aligned better with specialized work than rules created from a task prompt alone.

Why it mattersReal successful work gives you a stronger starting point than imagined standards.

Try thisCollect examples that experts agree are good work.

How to use it

Suggested steps

  1. Collect examples that experts agree are good work.
  2. Write the specific qualities that make them successful.
  3. Test those rules on strong and weak examples before reusing them.
Start with evidence of qualityConceptual explanation · Evals Weekly
CheckWhat to look for
Task descriptionWhat was requested
Successful workWhat meeting that request looks like

Keep in mind Professional roles have different requirements. A useful rule in one role may not transfer.

Open post
Source & why it’s here

FinProBench (opens in a new tab)

Original issue reading order, preserved from the archive.

Stop testing early only when the evidence supports it.

Published August 3, 2026

What it saysThe paper describes deciding whether to accept, reject, continue testing, or withhold a decision before every test has run. Some tasks reached a conclusion much sooner than others.

Why it mattersA small sample is useful only if it supports the decision you need to make.

Try thisSet the smallest improvement that would change your decision.

How to use it

Suggested steps

  1. Set the smallest improvement that would change your decision.
  2. Require coverage of important task types.
  3. Continue testing when the result is too uncertain to call.
Let uncertainty decideConceptual explanation · Evals Weekly
CheckWhat to look for
Evidence supports a clear decisionStop according to the rule
Evidence is still too weakContinue or withhold a decision

Keep in mind Early stopping requires valid decision rules. Do not stop simply because the first results look good.

Open post
Source & why it’s here

ParEvalLayer (opens in a new tab)

Original issue reading order, preserved from the archive.

Keep useful test examples, not every example.

Published August 2, 2026

What it saysThis work describes maintaining a test collection by the abilities its cases cover. New examples can be added, replace weaker ones, or go to human review.

Why it mattersA growing collection should improve coverage without filling up with duplicate tests.

Try thisTag each example with the ability and failure type it tests.

How to use it

Suggested steps

  1. Tag each example with the ability and failure type it tests.
  2. Check whether a new example adds coverage or improves an existing test.
  3. Keep the strongest representatives and review uncertain additions.
Give every test a jobConceptual explanation · Evals Weekly
CheckWhat to look for
New failure typeAdd coverage
Duplicate exampleKeep the more useful test

Keep in mind Keep important rare failures even if they do not occur often.

Open post
Source & why it’s here

Who Belongs in the Eval Set? (opens in a new tab)

Original issue reading order, preserved from the archive.

About this publication

Evals Weekly, by Kayzn.

We turn research and practical lessons about AI testing into a short read: what matters, what it means, and how to use it.

Kayzn is led by Shehaaz Saif and Elias Haroun. Shehaaz brings experience in AI quality at Expedia Group; Elias brings engineering experience from Amazon.

Better checks. Clearer decisions.

Kayzn is an AI evaluation consultancy. We help teams define a good answer, check whether AI scores agree with their experts, and keep testing as their agents change.

See how Kayzn can help

How we choose

Worth your five minutes.

We find useful ideas about testing AI, explain what the authors actually found, and suggest a small next step. Research papers, engineering posts, and public discussions all belong here. Dates stay visible, including when we revisit an older idea.

We favor clear evidence and practical value. Social attention helps us discover work; it does not prove the work is right.

Summaries are written with AI assistance and checked against available sources. Suggested steps are our application of the work.

How we rank the reading list

Each available signal gets a 0–5 rating. The weighted total sets the order within each new issue.

Relevance to testing AI30%
Strength of the evidence25%
Practical usefulness20%
A genuinely new idea10%
Recent publication5%
Attention from relevant experts7%
Public engagement3%

Unknown signals are left out and the remaining weights are rescaled. Expert attention requires an attributable public link. Engagement uses dated, platform-relative counts when available; likes alone never establish quality. A paper and a post about the same work count as one story.

For this issue, comparable engagement counts were unavailable, so the order reflects editorial value. Public mentions are linked where found. We are not claiming these are the most popular sources. Older issues retain their original reading order.