From One score cannot tell the whole story.Issue published Aug 14, 2026

The same model can behave very differently in a different setup.

Published August 9, 2026

What it saysChanging the tools and control logic around a coding model changed its cost and failure patterns, even when task success rates were fairly close.

Why it mattersA model score is also a score for the setup around it.

Try thisRecord the model, instructions, tools, retry limits, and stopping rules for each test.

How to use it

Suggested steps

  1. Record the model, instructions, tools, retry limits, and stopping rules for each test.
  2. Change one part at a time.
  3. Compare success, cost, and failure type together.
Make comparisons fairConceptual explanation · Evals Weekly
CheckWhat to look for
Same model, different toolsA different system is being tested
Same model and setupResults are easier to compare

Keep in mind Results from coding tasks may not transfer to other products.

A quick check for yourself

What is the practical lesson?

Show answer

A model score is also a score for the setup around it.

Source & why it’s here

The Scaffold Effect in Coding Agents (opens in a new tab)

Original issue reading order, preserved from the archive.