From Test the score before trusting it.Issue published Sep 18, 2026

When the data changes, recompute the answer.

Published Sep 15, 2026Read first

What it saysAdobe researchers compute reference answers with code at test time, then compare individual facts rather than wording. On synthetic data, this agreed with three expert reviewers better than written instructions for computing the answer.

Why it mattersA stored answer can become wrong when prices, availability, or customer records change. The checker needs the same data state as the agent.

Try thisPick one changing-data question. Write a small, independently reviewed calculation for its expected answer.

How to use it

Suggested steps

  1. Run the agent and reference calculation against the same data snapshot.
  2. Compare claimed facts and missing facts separately.
  3. Have experts check disagreements before trusting the score.
One data state, two independent pathsConceptual explanation · Evals Weekly
PathJob
AgentAnswer the user using current data
Reference codeCompute expected facts from that same data
CheckerFind wrong claims and missing facts

Keep in mind Only 53 comparable pass/fail cases, using synthetic data and the same model for agent and judge. Reference code can contain mistakes.

A quick check for yourself

What if the reference reads newer data than the agent?

Show answer

A timing mismatch can look like an agent error. Use the same data snapshot.

Source & why it’s here

Skill-based Agentic Evaluation for Real-time Data Science Tasks (opens in a new tab)

Read first: a concrete method checked against expert decisions. Extends the previous issue’s call to check actual results.