Test with synthetic customers before real ones.
Published Sep 24, 2026Read first
What it saysNubank compared agent versions through multi-turn synthetic customers and mocked tool responses. A configuration selected after more than 16,000 simulated conversations later improved self-service by 8.82 percentage points in a live test, with no significant change in customer satisfaction.
Why it mattersThis adds a useful layer between hand-written tests and exposing real customers. Simulation finds weak candidates cheaply; customer behavior still confirms the release.
Try thisChoose one failure hypothesis. Run the incumbent and candidate through the same 100 simulated journeys before a small live test.
How to use it
Suggested steps
- Keep user scenarios and mocked tool states identical for both versions.
- Calibrate each new checker with domain experts before using it to screen.
- Only advance clear wins; confirm them with a real product outcome.
| Layer | Decision it supports |
|---|---|
| Simulation | Is this candidate safe enough to test? |
| Expert calibration | Does the checker apply the intended rule? |
| Live outcome | Did the customer experience improve? |
Keep in mind The synthetic customers wrote much longer messages than real users, and mocked tools did not test backend behavior, latency, saved state, or side effects. Treat simulation as a screen, not proof.
A quick check for yourself
Can a simulated win approve a release by itself?
Show answer
No. It can screen candidates; a real product outcome should confirm the winner.
Source & why it’s here
Read first: the strongest bridge this week between offline evaluation and real customer outcomes.