Issue 01

3 min read

Better answers, not prettier words.

Checkers can be distracted by writing style. This issue also looks at learning from real work, testing efficiently, and choosing useful examples.

Polished wording can fool an AI checker.

Published August 3 · revised August 4, 2026Read first

What it saysWhen the same substance was written in different styles, judges changed their preferences. A style-aware checking step helped separate the content from its presentation.

Why it mattersAn answer should not get credit for sounding better when its facts are unchanged.

Try thisWrite two versions of an answer with the same facts but different styles.

How to use it

Suggested steps

  1. Write two versions of an answer with the same facts but different styles.
  2. Check whether the factual score changes.
  3. Separate the accuracy question from the writing-quality question.
One answer, two questionsConceptual explanation · Evals Weekly
CheckWhat to look for
Are its facts correct?Check the evidence
Is it easy to read?Check the writing

Keep in mind Style matters for readability; the point is to stop it distorting unrelated scores.

A quick check for yourself

What is the practical lesson?

Show answer

An answer should not get credit for sounding better when its facts are unchanged.

Open post
Source & why it’s here

Style Wins, Substance Loses (opens in a new tab)

Original issue reading order, preserved from the archive.

Write quality rules from real examples of good work.

Published August 4, 2026

What it saysThis study builds scoring rules from professional deliverables. Those rules aligned better with specialized work than rules created from a task prompt alone.

Why it mattersReal successful work gives you a stronger starting point than imagined standards.

Try thisCollect examples that experts agree are good work.

How to use it

Suggested steps

  1. Collect examples that experts agree are good work.
  2. Write the specific qualities that make them successful.
  3. Test those rules on strong and weak examples before reusing them.
Start with evidence of qualityConceptual explanation · Evals Weekly
CheckWhat to look for
Task descriptionWhat was requested
Successful workWhat meeting that request looks like

Keep in mind Professional roles have different requirements. A useful rule in one role may not transfer.

A quick check for yourself

What is the practical lesson?

Show answer

Real successful work gives you a stronger starting point than imagined standards.

Open post
Source & why it’s here

FinProBench (opens in a new tab)

Original issue reading order, preserved from the archive.

Stop testing early only when the evidence supports it.

Published August 3, 2026

What it saysThe paper describes deciding whether to accept, reject, continue testing, or withhold a decision before every test has run. Some tasks reached a conclusion much sooner than others.

Why it mattersA small sample is useful only if it supports the decision you need to make.

Try thisSet the smallest improvement that would change your decision.

How to use it

Suggested steps

  1. Set the smallest improvement that would change your decision.
  2. Require coverage of important task types.
  3. Continue testing when the result is too uncertain to call.
Let uncertainty decideConceptual explanation · Evals Weekly
CheckWhat to look for
Evidence supports a clear decisionStop according to the rule
Evidence is still too weakContinue or withhold a decision

Keep in mind Early stopping requires valid decision rules. Do not stop simply because the first results look good.

A quick check for yourself

What is the practical lesson?

Show answer

A small sample is useful only if it supports the decision you need to make.

Open post
Source & why it’s here

ParEvalLayer (opens in a new tab)

Original issue reading order, preserved from the archive.

Keep useful test examples, not every example.

Published August 2, 2026

What it saysThis work describes maintaining a test collection by the abilities its cases cover. New examples can be added, replace weaker ones, or go to human review.

Why it mattersA growing collection should improve coverage without filling up with duplicate tests.

Try thisTag each example with the ability and failure type it tests.

How to use it

Suggested steps

  1. Tag each example with the ability and failure type it tests.
  2. Check whether a new example adds coverage or improves an existing test.
  3. Keep the strongest representatives and review uncertain additions.
Give every test a jobConceptual explanation · Evals Weekly
CheckWhat to look for
New failure typeAdd coverage
Duplicate exampleKeep the more useful test

Keep in mind Keep important rare failures even if they do not occur often.

A quick check for yourself

What is the practical lesson?

Show answer

A growing collection should improve coverage without filling up with duplicate tests.

Open post
Source & why it’s here

Who Belongs in the Eval Set? (opens in a new tab)

Original issue reading order, preserved from the archive.

Explore another issue.

Back to the archive