Issue 02

3 min read

One score cannot tell the whole story.

A good average can hide missed mistakes, wasteful steps, and unreliable behavior. Start with the failures that matter to your users.

Choose a checker by the mistakes it catches.

Published August 9, 2026Read first

What it saysJudges with similar overall accuracy differed sharply in detecting unwanted side effects. Catching more problems also sent more work to human reviewers.

Why it mattersA good average can hide the very failures you need to catch.

Try thisList your most important failure types.

How to use it

Suggested steps

  1. List your most important failure types.
  2. Check how many real problems each judge catches and how many false alarms it creates.
  3. Choose a setting your review team can handle; recheck after changes.
Measure both costsConceptual explanation · Evals Weekly
CheckWhat to look for
Missed mistakesBad answers that passed
False alarmsGood answers sent for review

Keep in mind Higher detection can create more review work. The best trade-off depends on the product.

A quick check for yourself

What is the practical lesson?

Show answer

A good average can hide the very failures you need to catch.

Open post
Source & why it’s here

Selecting LLM Judges for Agent Evaluation Pipelines (opens in a new tab)

Original issue reading order, preserved from the archive.

A better-looking answer can come from a broken process.

Published August 9, 2026

What it saysIn analytics tasks, a stronger configuration answered more questions using real data while still showing unreliable choices during its work. The final answer alone missed these issues.

Why it mattersCheck how the assistant reached an answer as well as the answer itself.

Try thisSave the data and tool calls used for a task.

How to use it

Suggested steps

  1. Save the data and tool calls used for a task.
  2. Check that the assistant interpreted the right tables and respected its limits.
  3. Repeat the task to see whether the process stays reliable.
Check the answer and the pathConceptual explanation · Evals Weekly
CheckWhat to look for
Final responseIs the answer supported?
Work recordDid it use the right data and steps?

Keep in mind These observations come from an enterprise analytics setup, not every assistant.

A quick check for yourself

What is the practical lesson?

Show answer

Check how the assistant reached an answer as well as the answer itself.

Open post
Source & why it’s here

Evaluating Enterprise Analytics Agents (opens in a new tab)

Original issue reading order, preserved from the archive.

The same model can behave very differently in a different setup.

Published August 9, 2026

What it saysChanging the tools and control logic around a coding model changed its cost and failure patterns, even when task success rates were fairly close.

Why it mattersA model score is also a score for the setup around it.

Try thisRecord the model, instructions, tools, retry limits, and stopping rules for each test.

How to use it

Suggested steps

  1. Record the model, instructions, tools, retry limits, and stopping rules for each test.
  2. Change one part at a time.
  3. Compare success, cost, and failure type together.
Make comparisons fairConceptual explanation · Evals Weekly
CheckWhat to look for
Same model, different toolsA different system is being tested
Same model and setupResults are easier to compare

Keep in mind Results from coding tasks may not transfer to other products.

A quick check for yourself

What is the practical lesson?

Show answer

A model score is also a score for the setup around it.

Open post
Source & why it’s here

The Scaffold Effect in Coding Agents (opens in a new tab)

Original issue reading order, preserved from the archive.

Before using a leaderboard, ask what it really tests.

Published August 9, 2026

What it saysModels changed position across different agent benchmarks. A broad score can mix reasoning, tool use, and other abilities in ways that hide what improved.

Why it mattersChoose tests that reflect the work your users need done.

Try thisName the specific ability you need to measure.

How to use it

Suggested steps

  1. Name the specific ability you need to measure.
  2. Check which tasks in a benchmark actually require that ability.
  3. Test critical abilities separately before combining scores.
Start with the intended useConceptual explanation · Evals Weekly
CheckWhat to look for
Your user’s taskThe ability that matters
Benchmark taskEvidence it measures that ability

Keep in mind Different rankings are not automatically an error: different tests may measure different abilities.

A quick check for yourself

What is the practical lesson?

Show answer

Choose tests that reflect the work your users need done.

Open post
Source & why it’s here

Construct Validity Failures in Agentic AI Benchmarks (opens in a new tab)

Original issue reading order, preserved from the archive.

Explore another issue.

Back to the archive