Issue 08

3 min read

Screen broadly. Confirm with real outcomes.

Three new studies improve the path from an offline score to a trustworthy release decision. Start with Nubank’s production-tested approach: simulate customer journeys, then let real customer outcomes decide what ships.

Test with synthetic customers before real ones.

Published Sep 24, 2026Read first

What it saysNubank compared agent versions through multi-turn synthetic customers and mocked tool responses. A configuration selected after more than 16,000 simulated conversations later improved self-service by 8.82 percentage points in a live test, with no significant change in customer satisfaction.

Why it mattersThis adds a useful layer between hand-written tests and exposing real customers. Simulation finds weak candidates cheaply; customer behavior still confirms the release.

Try thisChoose one failure hypothesis. Run the incumbent and candidate through the same 100 simulated journeys before a small live test.

How to use it

Suggested steps

  1. Keep user scenarios and mocked tool states identical for both versions.
  2. Calibrate each new checker with domain experts before using it to screen.
  3. Only advance clear wins; confirm them with a real product outcome.
Conceptual release ladderConceptual explanation · Evals Weekly
LayerDecision it supports
SimulationIs this candidate safe enough to test?
Expert calibrationDoes the checker apply the intended rule?
Live outcomeDid the customer experience improve?

Keep in mind The synthetic customers wrote much longer messages than real users, and mocked tools did not test backend behavior, latency, saved state, or side effects. Treat simulation as a screen, not proof.

A quick check for yourself

Can a simulated win approve a release by itself?

Show answer

No. It can screen candidates; a real product outcome should confirm the winner.

Open post
Source & why it’s here

Ten judges may share one blind spot.

Published Sep 18, 2026

What it saysIn the main ten-judge test, shared errors reduced the independent information to roughly 3.5 judges. In up to 28% of system comparisons, a claimed significant difference disappeared after the analysis accounted for those shared mistakes.

Why it mattersAdding more judges can create false confidence when they fail on the same examples. Using models from different providers does not guarantee independent evidence.

Try thisRun every judge on about 100 trusted examples. Look for cases that fool several judges before choosing a voting rule.

How to use it

Suggested steps

  1. Use the same trusted set and hidden labels for every judge.
  2. Group shared mistakes by failure type, not only by model.
  3. Choose the voting rule on one set; verify it on held-out examples.
Conceptual view of shared errorsConceptual explanation · Evals Weekly
What you seeWhat may be true
8 of 10 judges agreeSeveral copied the same blind spot
Trusted examples expose clustersAgreement can be weighted more honestly

Keep in mind The main bank used five small open models with two prompt styles. Controlled tests focused on position bias, and the numerical results should not be assumed to transfer to another judge bank.

A quick check for yourself

Does using different model providers make judge votes independent?

Show answer

Not necessarily. Check shared mistakes on trusted examples.

Open post
Source & why it’s here

Send uncertain scores to a stronger checker.

Published Sep 22, 2026

What it saysA cheap first-pass judge accepted confident decisions and sent uncertain ones to a stronger judge. On pooled public tasks, one threshold escalated 34% of cases and reached 91.3% accuracy versus 91.7% for always using the stronger judge, at 47% of its fee.

Why it mattersA calibrated escalation rule can reduce evaluation cost without treating every case as equally easy. Confidence is useful for routing, not proof that a verdict is right.

Try thisTest a confidence threshold on local human-labeled examples. Track both accuracy and how many cases require escalation.

How to use it

Suggested steps

  1. Choose the threshold on local examples and freeze it before testing.
  2. Escalate invalid outputs, uncertain cases, and critical decisions.
  3. Report error recall, false alarms, and review load for each task group.
Conceptual escalation ruleConceptual explanation · Evals Weekly
Observed caseRoute
High confidence on a validated taskAccept first-pass verdict
Low confidence or critical decisionUse stronger check or human review
Known weak task or styleBypass the first pass

Keep in mind The first-pass judge is proprietary. Human adjudication covered 183 disagreement-selected items, used one author as annotator, and left 36 labels undecided. The threshold failed to transfer to every task.

A quick check for yourself

Is high confidence the same as correctness?

Show answer

No. It is a routing signal that must be checked on your own tasks.

Open post
Source & why it’s here

Explore another issue.

Back to the archive