Alp Cetin

San Francisco

eval-power

How many benchmark items does it take to tell two language models apart?

Question

Leaderboards rank models by accuracy differences that are often a point or less. Two questions follow. How many test items does a comparison need before a gap that small is real? And if you plan that number from a small pilot, as people do, how often does the plan deliver the power it promised?

Method

The data are per-question right and wrong answers for 395 models from the 2024 Open LLM Leaderboard on five benchmarks, about 28,000 items in all. Because both models answer the same questions, the comparison can be paired, which removes the difficulty the two models share. To test pilot planning, I split items into a small pilot and a held-out pool, sized a full study from the gap seen in the pilot, and checked how often that study detected the difference on held-out items.

Result

Most leaderboard gaps are noise. Even without correcting for multiple comparisons, only 1 to 12 of the 394 neighboring pairs in each benchmark’s ranking differ significantly. Detecting a one-point gap at 80% power takes roughly 6,000 to 17,000 items per model when the comparison is paired, and 22,000 to 38,000 when it isn’t; most of these benchmarks have far fewer items than that.

Detection rate delivered by pilot-based plans that target 80% powerFor each benchmark, plans sized from the gap seen in a 128-item or 256-item pilot were checked on held-out items. Median detection ranged from 49% to 81%, mostly below the 80% target, with wide spread between model pairs.pilot of 128 itemspilot of 256 itemsdots: median over model pairs; bars: middle half0%20%40%60%80%100%80% targetdetection rate on held-out itemsARC-ChallengeGSM8KWinoGrandeHellaSwagMMLU
How often a study planned from a pilot’s observed gap actually detected the difference on held-out items, for 40 model pairs per benchmark. Dots are medians, bars the middle half of pairs. Data: calibration.csv.

Pilots are a poor guide. Studies sized from a 128-item pilot to have 80% power detected the difference only 49% to 79% of the time, depending on the benchmark, and a larger pilot didn’t reliably help. The gap a small pilot sees is itself noisy, so planning around it overshoots and undershoots by wide margins. This is a known problem in statistics; the safer plan is to fix the smallest difference that matters before collecting any data.

Prospective study

I then generated 13,830 primary-study answers from six small models: Qwen2.5 at 1.5B, 3B and 7B; SmolLM2-1.7B; Phi-3.5-mini; and Mistral-7B-v0.3. Each answered the same 64 pilot questions per benchmark once greedily and five times at temperature 0.7. I committed the sample-size plans before collecting fresh-item answers. The protocol pins models and decoding; the study report records the answer counts.

Twenty-three of 27 planned comparisons detected a difference: 13/14 on GSM8K and 10/13 on guided direct-choice ARC-Challenge. Each feasible pair used its planned 16–264 fresh questions and a paired t-test on five-answer item means. This is an observed fraction, not proof of 80% power: effects differ, and pairs share models and answers. Three plans exceeded the declared 512-question pool and were not run; I did not cap them and call them 80%-power tests.

Decoding contributes a much larger share of variance on GSM8K. At five stochastic answers per question, its median estimated share is 33.2%, against 4.6% on direct-choice ARC. I estimate variance within questions, subtract its contribution from the observed variance of paired item means, and retain the unclipped estimate before applying a zero floor for planning. These are finite-pilot estimates, not perfectly identified population components.

Decoding-noise share differs sharply between GSM8K and direct-choice ARCEstimated decoding contribution to paired item-mean variance with five stochastic answers per question. Fifteen model pairs per benchmark; median share is 33.2% on GSM8K and 4.6% on guided direct-choice ARC. These finite-pilot estimates are descriptive, not known population variance components.five stochastic answers/question; 15 model pairs/benchmarksmall dots: pairs; large dots: median0%20%40%60%GSM8K33.2%ARC (direct choice)4.6%estimated decoding share of total variance at k=5
Estimated decoding contribution to paired item-mean variance at five stochastic answers per question. Small dots are the 15 model pairs per benchmark; large dots are medians. Pairs share models and questions, so these are not independent replicates. Data: summary.json.

This confirms plans fixed after a pilot, not a protocol registered before seeing any pilot. I fixed numeric grading and a random-seed reuse bug, recollected the primary pilots, and changed all six ARC models to guided Answer: <label> generation before confirmation. Direct choice is not conventional ARC option-likelihood evaluation. The revision comparison retains those changes.

Historical votes

In a separate historical Arena release, 54,985 retained votes cover 55 models. Only four of 54 adjacent model contrasts exclude zero under either pointwise Wald or bootstrap 95% intervals. These are historical votes, not newly collected preferences or a current leaderboard, and the adjacent intervals are not simultaneous or rank-selection adjusted. Data: Arena summary.

Limits

The earlier leaderboard analysis uses older open models, many of them fine-tunes or merges, with each answer recorded once. Its resampling of held-out answers is separate from the new prospective collection. Neither study is a current frontier-model benchmark.

The prospective study uses public benchmarks and small instruction-tuned models; contamination is possible. A finite pilot can misestimate both effects and variance. Comparisons overlap and have no multiplicity correction. Five stochastic draws are not greedy decoding, and fewer questions do not imply fewer generated answers or lower compute cost. The Arena model assumes IID votes and transitive preference strength; retained aggregates cannot recover unavailable user or prompt clustering.