From The right score needs the right evidence.Issue published Aug 21, 2026

Finishing a task is only the first check.

Published August 19, 2026

What it saysThis benchmark follows assistants through hundreds of decisions in a long simulation. All tested models finished, yet the quality of their outcomes differed and became clearer late in the task.

Why it mattersA completion count can hide weak decisions along the way.

Try thisChoose one workflow with several linked decisions.

How to use it

Suggested steps

  1. Choose one workflow with several linked decisions.
  2. Check intermediate decisions as well as the final outcome.
  3. Repeat long enough to see whether early mistakes accumulate.
Look beyond completionConceptual explanation · Evals Weekly
CheckWhat to look for
Did it finish?Basic completion
Did its decisions work out?Quality over the whole task

Keep in mind A simulated environment cannot establish how the same behavior will perform in the real world.

A quick check for yourself

What is the practical lesson?

Show answer

A completion count can hide weak decisions along the way.

Source & why it’s here

FM-Bench (opens in a new tab)

Original issue reading order, preserved from the archive.