Issue 05

3 min read

Reuse the work. Verify the result.

Three ways to make AI testing more useful: reuse earlier reviews, check changing facts, and verify exactly what an assistant changed.

Stop paying to check the same evidence twice.

Published September 2, 2026Read first

What it saysWhen comparing systems that find documents, many results overlap. This method saves earlier judgments and reviews only new documents, then compares every system using the updated shared collection.

Why it mattersEach round of testing can make the next round cheaper.

Try thisCompare the documents returned by two search versions.

How to use it

Suggested steps

  1. Compare the documents returned by two search versions.
  2. Keep earlier judgments when the document and scoring rule are unchanged.
  3. Judge the new documents; compare both versions on the shared collection.
What needs a fresh review?Conceptual explanation · Evals Weekly
CheckWhat to look for
Same document, same ruleReuse the judgment
New document or changed ruleReview again

Keep in mind Recheck judgments when the content or quality rule changes. Report uncertainty around close results.

A quick check for yourself

What is the practical lesson?

Show answer

Each round of testing can make the next round cheaper.

Open post
Source & why it’s here

Incremental Pooled LLM Evaluation (opens in a new tab)

Original issue reading order, preserved from the archive.

When the facts change, does the AI change its plan?

Published September 3, 2026

What it saysThis test gives assistants conflicting information from users, stored knowledge, and tools. The tested models did not reliably handle every kind of conflict.

Why it mattersMentioning a new fact is not enough. The assistant must use it in its next action.

Try thisCreate a test where an old availability result conflicts with a fresh one.

How to use it

Suggested steps

  1. Create a test where an old availability result conflicts with a fresh one.
  2. Check which result the assistant uses when it acts.
  3. Inspect the resulting record and review any mistaken decisions.
A changed fact should change the actionConceptual explanation · Evals Weekly
CheckWhat to look for
Old result: availableAssistant has an initial plan
Fresh result: unavailablePlan and action must update

Keep in mind A controlled test does not cover every conflict a real product will encounter.

A quick check for yourself

What is the practical lesson?

Show answer

Mentioning a new fact is not enough. The assistant must use it in its next action.

Open post
Source & why it’s here

KC-Bench (opens in a new tab)

Original issue reading order, preserved from the archive.

Let AI propose the change. Make code apply it precisely.

Published August 31, 2026

What it saysIn a study of editing server configuration files, apparently successful edits could change the wrong thing. The proposed approach separates deciding what to change from applying the exact edit.

Why it mattersA correct request can still fail during execution.

Try thisChoose one action where the assistant changes a saved record.

How to use it

Suggested steps

  1. Choose one action where the assistant changes a saved record.
  2. Have it specify the record, field, and new value.
  3. Use code to validate and apply that change; stop if the target is unclear.
Two jobs, two checksConceptual explanation · Evals Weekly
CheckWhat to look for
Decide what should changeReview the proposed action
Apply the changeCheck the exact saved result

Keep in mind The study concerns a specific file-editing task. The proposed change still needs to be appropriate.

A quick check for yourself

What is the practical lesson?

Show answer

A correct request can still fail during execution.

Open post
Source & why it’s here

Don’t Let the Model Write the YAML (opens in a new tab)

Original issue reading order, preserved from the archive.

Explore another issue.

Back to the archive