- Published on
Eval Scores Move Between Runs: How to Compare Two Versions Honestly
- Authors

- Name
- Mehdi Akiki
Article · Measurement
If you want to know whether version B is better than version A, run both on the same items and report an interval on the per-item difference. Do not compare two separate scores against a fixed threshold. In the simulation below, a gate that fails the build when the score drops more than two points below the recorded baseline goes red 28.6% of the time on a 100 item suite where absolutely nothing changed.
I have watched a team disable an eval gate for this reason. It produced red builds nobody could explain, and everyone learned to rerun it until it passed.
A simulation model of eval run to run noise
I simulate a pool of 4000 items, each with a latent difficulty. A version has a competence threshold and passes an item with a probability that is high for easy items, low for hard ones, and uncertain in a narrow band between them.
This shape matters. If every item were a coin flip, pairing would buy nothing. In a real eval set most items almost always pass or almost always fail, and only a minority is flaky.
The thresholds are solved numerically so version A scores exactly 70.0% and version B exactly 75.0%. B is 5.0 points better by construction. Everything below asks how easy that real difference is to see.
node experiments/ai-evals/compare-versions.mjs
node --test experiments/ai-evals/*.test.mjs
pool: 4000 items, true mean A 70.00%, true mean B 75.00%, true gain 5.0 pts
A fixed threshold gate flakes with no change at all
The suite is fixed, as in a repository. The first run records the baseline, later runs change nothing, and the gate fails when a run scores more than 2.0 points below it.
Each row below draws its own suite from the pool, so the centre of each band moves with the difficulty of that draw. Compare the widths across rows, not the levels.
[1] a fixed threshold gate on an UNCHANGED version A
n= 50 red builds 29.2% run to run 5th-95th percentile 72.0% to 84.0% (spread 12.0 pts)
n= 100 red builds 28.6% run to run 5th-95th percentile 63.0% to 74.0% (spread 11.0 pts)
n= 300 red builds 19.5% run to run 5th-95th percentile 64.3% to 70.0% (spread 5.7 pts)
n=1000 red builds 6.8% run to run 5th-95th percentile 68.5% to 71.6% (spread 3.1 pts)
Almost three builds in ten go red on the 100 item suite with no code, prompt, or model change. Part of that noise is the run, and part is the baseline, itself recorded on one ordinary run.
At 1000 items the gate is defensible. At 50 or 100 items it reports noise.
How wide is the interval on a single score
The second experiment puts a percentile bootstrap interval around one measured score.
[2] bootstrap 95% CI width for a single measured score
n= 50 score 66.0% CI [52.0%, 78.0%] width 26.0 pts
n= 100 score 62.0% CI [53.0%, 71.0%] width 18.0 pts
n= 300 score 69.3% CI [64.3%, 74.3%] width 10.0 pts
n=1000 score 69.2% CI [66.3%, 72.0%] width 5.7 pts
The width falls roughly with the inverse square root of the item count. To halve the interval you need four times the items. Each row is again a separate draw, which is why the scores wander around the true 70.0%.
So a single score from a 50 item suite carries a 26 point interval, and a 5 point improvement is invisible inside it. This is why I report the interval next to the accuracy in What Evidence an AI Feature Needs Before You Ship It, where a different fixture lands on a very similar 10.1 point width at 300 items.
Why a paired comparison beats an unpaired one
I run 4000 simulated experiments per configuration and count how often each method declares B better at the 5% level.
Paired means both versions see the same items and the test runs on the per-item difference. Unpaired means each version gets its own independent sample.
[3] power to detect the real 5.0 point gain (alpha = 0.05)
paired = same items for both versions, one run each
Monte Carlo standard error at 4000 trials: 0.8 pts near 50%, 0.3 pts near 95%
n= 50 paired 12.3% McNemar 7.2% unpaired 8.1%
n= 100 paired 22.0% McNemar 13.5% unpaired 12.6%
n= 300 paired 51.8% McNemar 46.8% unpaired 27.5%
n=1000 paired 95.9% McNemar 95.4% unpaired 69.9%
At 300 items the paired test finds the real gain 51.8% of the time, the unpaired test 27.5%. At 1000 items the gap is 95.9% against 69.9%.
Read those figures against the Monte Carlo error the script prints. 95.9% against 95.4% is not a difference. 51.8% against 46.8% is. The paired test also uses a normal approximation rather than a t distribution, so the 50 item row is a small overstatement.
The reason is simple. In an unpaired comparison most of the variance comes from the item sample: version B might have drawn easier items. Pairing removes that source completely, because both versions face the same difficulty. What remains is the run to run flakiness on the uncertain items.
McNemar looks only at the discordant items. It tracks the paired test closely and is easier to explain to a reviewer.
The top row is the uncomfortable one. Even paired, a 50 item suite finds a real 5 point gain 12.3% of the time. A suite that size cannot resolve five points. It can only resolve differences of roughly twenty points.
Are repeated runs per item worth it
The next idea is to run each item several times and average. It helps, but the accounting matters.
At a fixed budget of 600 runs per version:
[4] repeated runs per item at a fixed budget of 600 executions per version
items 600 x 1 runs = 600 executions power 81.3%
items 300 x 2 runs = 600 executions power 82.2%
items 200 x 3 runs = 600 executions power 80.2%
items 120 x 5 runs = 600 executions power 79.9%
items 60 x 10 runs = 600 executions power 77.0%
With a Monte Carlo error near 0.6 points, the first three plans are the same result. The curve is flat up to two or three repeats and then bends down: 60 items run 10 times each loses about 5 points against 300 items run twice. So repeats past three buy nothing at a fixed budget, and past that they cost.
But if the item count is already fixed by how many cases you could label, repeats do buy a lot:
[5] repeated runs per item at a fixed item count of 100
items 100 x 1 runs = 100 executions power 21.6%
items 100 x 3 runs = 300 executions power 50.4%
items 100 x 5 runs = 500 executions power 71.7%
items 100 x 10 runs = 1000 executions power 93.7%
My reading: spend on labelling new items first, and on repeats second, and never take a plan past three repeats when the budget is what limits you. Repeats are the tool when the labelled set cannot grow this week. They also expose instability on a single case, which is a different purpose, described in "Golden Tests for Non-Deterministic AI Outputs".
A recipe for comparing two versions honestly
First, freeze the item set before the comparison. If items are added between the baseline and the candidate, the comparison is unpaired again whatever the code says.
Second, run both versions on that set in the same session, with the same retrieval snapshot and grader version. Changing the judge between runs makes the difference uninterpretable, which is why the judge gets a frozen version, as in "Calibrating an LLM Judge Against Human Disagreement".
Third, keep the per-item outcomes, not only the totals. Without them there is no pairing and no way to see which items flipped.
Fourth, report the difference with its interval, and state the direction. "Plus 3.7 points, interval minus 3.3 to plus 10.3", which is what the triage fixture in the pillar article reports, is an honest sentence. "Plus 3.7 points" alone is not.
Fifth, keep safety cases outside this arithmetic. A trajectory violation is not a statistical question and does not get averaged against a quality gain. "Regression Testing Across Prompt and Model Changes" covers how to isolate which layer moved once the comparison is sound.
What I check before I trust an eval comparison
- Both versions ran on exactly the same item set, in the same session.
- Per-item results are stored, so the comparison can be paired afterwards.
- The reported number is a difference with an interval, not two separate scores.
- The item count is large enough for the size of the effect being claimed.
- The grader, the retrieval snapshot, and the tool versions did not change between the two runs.
- Safety and policy cases are gated separately, at zero, and are not averaged.
- The gate threshold was written down before anyone saw the candidate's score.
Sources
- Card, Henderson, Khandelwal, Jia, Mahowald and Jurafsky, With Little Power Comes Great Responsibility
- Dror, Baumer, Shlomov and Reichart, The Hitchhiker's Guide to Testing Statistical Significance in NLP
- McNemar, Note on the sampling error of the difference between correlated proportions or percentages
- Efron, Bootstrap Methods: Another Look at the Jackknife
- Brown, Cai and DasGupta, Interval Estimation for a Binomial Proportion
- Anthropic, Demystifying evals for AI agents