When the data changes, recompute the answer.
Published Sep 15, 2026Read first
What it saysAdobe researchers compute reference answers with code at test time, then compare individual facts rather than wording. On synthetic data, this agreed with three expert reviewers better than written instructions for computing the answer.
Why it mattersA stored answer can become wrong when prices, availability, or customer records change. The checker needs the same data state as the agent.
Try thisPick one changing-data question. Write a small, independently reviewed calculation for its expected answer.
How to use it
Suggested steps
- Run the agent and reference calculation against the same data snapshot.
- Compare claimed facts and missing facts separately.
- Have experts check disagreements before trusting the score.
| Path | Job |
|---|---|
| Agent | Answer the user using current data |
| Reference code | Compute expected facts from that same data |
| Checker | Find wrong claims and missing facts |
Keep in mind Only 53 comparable pass/fail cases, using synthetic data and the same model for agent and judge. Reference code can contain mistakes.
A quick check for yourself
What if the reference reads newer data than the agent?
Show answer
A timing mismatch can look like an agent error. Use the same data snapshot.
Source & why it’s here
Skill-based Agentic Evaluation for Real-time Data Science Tasks (opens in a new tab)
Read first: a concrete method checked against expert decisions. Extends the previous issue’s call to check actual results.