Check the result, not just the reply.
Published Jan 9, 2026
What it saysAn AI assistant can say it completed a task without completing it. Anthropic recommends combining software checks, AI judgments, and human review, chosen for the task.
Why it mattersA convincing “done” message tells you little about whether a booking, file edit, or support request succeeded.
Try thisFor one task, define the evidence that would prove it worked. Check that evidence directly.
How to use it
Suggested steps
- Turn a real failure into a repeatable test.
- Check the saved result with code where possible.
- Use people to check subjective quality; repeat the test after changes.
| Question | Useful check |
|---|---|
| Was the change saved? | Inspect the saved record |
| Was the answer useful? | Compare with human review |
Keep in mind Practice guidance; each product still needs its own success criteria.
A quick check for yourself
Is “I booked it” proof of a booking?
Show answer
No. Inspect the booking record.
Source & why it’s here
Demystifying evals for AI agents (opens in a new tab)
A practical starting point for testing an entire task.