Issue 03

3 min read

The right score needs the right evidence.

File versions, scoring scales, and longer tasks can all change the result. Four papers show what a single overall score leaves out.

Make sure the AI and its checker see the same file.

Published August 18, 2026Read first

What it saysThis work makes file state explicit. What an assistant can read, edit, and submit affects its measured performance; different representations can otherwise refer to different versions.

Why it mattersA version mismatch can look like an AI mistake or hide one.

Try thisRecord which file version the assistant reads and edits.

How to use it

Suggested steps

  1. Record which file version the assistant reads and edits.
  2. Link the readable text and original file to the same version.
  3. Check the submitted file, including the changes actually saved.
Keep the evidence togetherConceptual explanation · Evals Weekly
CheckWhat to look for
What the assistant readsThe correct source version
What the checker reviewsThe actual submitted version

Keep in mind Better file access can improve performance without changing the underlying model.

A quick check for yourself

What is the practical lesson?

Show answer

A version mismatch can look like an AI mistake or hide one.

Open post
Source & why it’s here

StagedWorkspace (opens in a new tab)

Original issue reading order, preserved from the archive.

Getting the order right does not make the scores right.

Published August 18, 2026

What it saysAutomatic scoring ranked the compared systems in the same order as the human jury, but compressed the differences between weak and strong results.

Why it mattersA score can pick the right winner and still set the wrong pass mark.

Try thisCompare AI and human scores for weak, average, and strong answers.

How to use it

Suggested steps

  1. Compare AI and human scores for weak, average, and strong answers.
  2. Check the size of the gaps as well as the order.
  3. Review examples near your release pass mark with people.
Two different questionsConceptual explanation · Evals Weekly
CheckWhat to look for
Who performed better?Compare the ranking
How much better?Compare the scoring scale

Keep in mind The study uses competition submissions. Your product needs its own human comparison.

A quick check for yourself

What is the practical lesson?

Show answer

A score can pick the right winner and still set the wrong pass mark.

Open post
Source & why it’s here

The IOL-AI Challenge (opens in a new tab)

Original issue reading order, preserved from the archive.

Finishing a task is only the first check.

Published August 19, 2026

What it saysThis benchmark follows assistants through hundreds of decisions in a long simulation. All tested models finished, yet the quality of their outcomes differed and became clearer late in the task.

Why it mattersA completion count can hide weak decisions along the way.

Try thisChoose one workflow with several linked decisions.

How to use it

Suggested steps

  1. Choose one workflow with several linked decisions.
  2. Check intermediate decisions as well as the final outcome.
  3. Repeat long enough to see whether early mistakes accumulate.
Look beyond completionConceptual explanation · Evals Weekly
CheckWhat to look for
Did it finish?Basic completion
Did its decisions work out?Quality over the whole task

Keep in mind A simulated environment cannot establish how the same behavior will perform in the real world.

A quick check for yourself

What is the practical lesson?

Show answer

A completion count can hide weak decisions along the way.

Open post
Source & why it’s here

FM-Bench (opens in a new tab)

Original issue reading order, preserved from the archive.

A better total can hide a worse trade-off.

Published August 15, 2026

What it saysThis study compares an assistant with professional estimators on real building plans. It checks how much material was identified separately from how accurate the quantities were.

Why it mattersFinding more items and measuring them correctly are different abilities.

Try thisSplit your overall score into the qualities that matter.

How to use it

Suggested steps

  1. Split your overall score into the qualities that matter.
  2. Compare the assistant with expert-reviewed examples on each quality.
  3. Inspect whether a higher total comes with an unacceptable weakness.
Show both sides of the resultConceptual explanation · Evals Weekly
CheckWhat to look for
CoverageWere the required items found?
AccuracyWere their quantities right?

Keep in mind The findings concern a specific estimation task and expert reference set.

A quick check for yourself

What is the practical lesson?

Show answer

Finding more items and measuring them correctly are different abilities.

Open post
Source & why it’s here

Handoff-H1 (opens in a new tab)

Original issue reading order, preserved from the archive.

Explore another issue.

Back to the archive