From Who checks the AI checker?Issue published Sep 13, 2026

Check the result, not just the reply.

Published Jan 9, 2026

What it saysAn AI assistant can say it completed a task without completing it. Anthropic recommends combining software checks, AI judgments, and human review, chosen for the task.

Why it mattersA convincing “done” message tells you little about whether a booking, file edit, or support request succeeded.

Try thisFor one task, define the evidence that would prove it worked. Check that evidence directly.

How to use it

Suggested steps

  1. Turn a real failure into a repeatable test.
  2. Check the saved result with code where possible.
  3. Use people to check subjective quality; repeat the test after changes.
Match the check to the questionConceptual explanation · Evals Weekly
QuestionUseful check
Was the change saved?Inspect the saved record
Was the answer useful?Compare with human review

Keep in mind Practice guidance; each product still needs its own success criteria.

A quick check for yourself

Is “I booked it” proof of a booking?

Show answer

No. Inspect the booking record.

Source & why it’s here

Demystifying evals for AI agents (opens in a new tab)

A practical starting point for testing an entire task.