From Screen broadly. Confirm with real outcomes.Issue published Sep 25, 2026

Test with synthetic customers before real ones.

Published Sep 24, 2026Read first

What it saysNubank compared agent versions through multi-turn synthetic customers and mocked tool responses. A configuration selected after more than 16,000 simulated conversations later improved self-service by 8.82 percentage points in a live test, with no significant change in customer satisfaction.

Why it mattersThis adds a useful layer between hand-written tests and exposing real customers. Simulation finds weak candidates cheaply; customer behavior still confirms the release.

Try thisChoose one failure hypothesis. Run the incumbent and candidate through the same 100 simulated journeys before a small live test.

How to use it

Suggested steps

  1. Keep user scenarios and mocked tool states identical for both versions.
  2. Calibrate each new checker with domain experts before using it to screen.
  3. Only advance clear wins; confirm them with a real product outcome.
Conceptual release ladderConceptual explanation · Evals Weekly
LayerDecision it supports
SimulationIs this candidate safe enough to test?
Expert calibrationDoes the checker apply the intended rule?
Live outcomeDid the customer experience improve?

Keep in mind The synthetic customers wrote much longer messages than real users, and mocked tools did not test backend behavior, latency, saved state, or side effects. Treat simulation as a screen, not proof.

A quick check for yourself

Can a simulated win approve a release by itself?

Show answer

No. It can screen candidates; a real product outcome should confirm the winner.

Source & why it’s here