Issue 06

3 min read

Who checks the AI checker?

Netflix’s latest post shows why AI checks need regular human review. Two earlier reads explain how to agree on the rules and test what actually happened.

An AI checker needs checkups, too.

Published Sep 4, 2026Read first

What it saysNetflix uses AI to check explanations of show recommendations. Experts supply examples and reasons; the team adjusts the checker, then keeps comparing its decisions with fresh human reviews.

Why it mattersA checker that worked at launch can become unreliable as the content changes. Matching a human’s verdict is not enough if the reason is wrong.

Try thisReview a small sample of recent answers. Record where people and the checker disagree, including why.

How to use it

Suggested steps

  1. Write one clear pass/fail rule with examples.
  2. Compare human decisions and reasons with the checker’s.
  3. Fix disagreements; retest on examples kept aside. Repeat regularly.
Keep the check connected to peopleConceptual explanation · Evals Weekly
WhenWhat to check
Before launchDoes it agree with experts, for the right reason?
After launchDoes that still hold on fresh examples?

Keep in mind A Netflix case study; results may differ for other products.

A quick check for yourself

When is the checker finished?

Show answer

It needs fresh human checks throughout its use.

Open post
Source & why it’s here

Agree on “good” before asking AI to score it.

Published Apr 29, 2026

What it saysPeople can use the same scoring rule and mean different things. MultEval lets a team compare disagreements, revise rules together, and keep the examples behind each decision.

Why it mattersAn AI checker cannot resolve a quality standard your team has never agreed on.

Try thisAsk two colleagues to review the same answers separately. Start the discussion with their disagreements.

How to use it

Suggested steps

  1. Choose examples that split opinion.
  2. Ask what each person understood the rule to mean.
  3. Rewrite the rule together; save a clear example beside it.
Find the ambiguityConceptual explanation · Evals Weekly
Vague ruleA clearer question
“Be helpful”Did the answer resolve the user’s request?
“Be accurate”Which claim does the source support?

Keep in mind A study of collaboration and a prototype, not proof of a universal accuracy gain.

A quick check for yourself

What does reviewer disagreement reveal?

Show answer

A rule may need clarification, not just a different AI model.

Open post
Source & why it’s here

Check the result, not just the reply.

Published Jan 9, 2026

What it saysAn AI assistant can say it completed a task without completing it. Anthropic recommends combining software checks, AI judgments, and human review, chosen for the task.

Why it mattersA convincing “done” message tells you little about whether a booking, file edit, or support request succeeded.

Try thisFor one task, define the evidence that would prove it worked. Check that evidence directly.

How to use it

Suggested steps

  1. Turn a real failure into a repeatable test.
  2. Check the saved result with code where possible.
  3. Use people to check subjective quality; repeat the test after changes.
Match the check to the questionConceptual explanation · Evals Weekly
QuestionUseful check
Was the change saved?Inspect the saved record
Was the answer useful?Compare with human review

Keep in mind Practice guidance; each product still needs its own success criteria.

A quick check for yourself

Is “I booked it” proof of a booking?

Show answer

No. Inspect the booking record.

Open post
Source & why it’s here

Demystifying evals for AI agents (opens in a new tab)

A practical starting point for testing an entire task.

Explore another issue.

Back to the archive