Accuracy comparison · item-budget planning

How many test items do I need?

I use a two-sided Gaussian approximation to plan a comparison between two models. Choose a meaningful gap before evaluation; an observed pilot gap can give an unreliable budget.

1. Plan a comparison

Accuracies are binary correct/incorrect rates. For unpaired studies, the budget is per model; for paired studies, both models answer the same items.

Target difference

Gap-only planning still needs a baseline to determine binary variance. A 2-point gap means 70% versus 72%, not a 2% relative increase.

Study design

Discordance counts items on which exactly one model is correct. Not every correlation is possible for two binary accuracies.

Error rates

Alpha is the false-positive rate for one comparison. Power is the chance of detecting the assumed gap if that gap and variance hold.

Items needed

—

per model

Choose the comparison, then calculate.


An approximation, not a guarantee. I assume independent, identically distributed items and a fixed variance. This is not exact McNemar power, a clustered-item budget, a finite-pool calculation, or a multiple-comparison correction.

For very small budgets or rare discordance, the Gaussian approximation can be poor. Keep the meaningful gap fixed; do not use a noisy pilot gap as a promise.

2. Did pilot plans deliver 80% power?

I checked plans from small pilots against disjoint heldout items in five historical benchmarks. This panel is a fixed result from committed data, not a prediction for the inputs above.

Loading committed results…

Each benchmark has 40 score-independent model pairs × 5 item splits. Each evaluable plan has 1,000 exact two-sided McNemar trials from the empirical heldout distribution, sampled with replacement. An estimated detection rate is not the true power; Monte Carlo bounds show simulation uncertainty only.

These are selected 395-model, 2024 Open LLM Leaderboard data, not current frontier models. Resampling existing items does not establish power on newly collected questions. The 80% target uses alpha 0.05, with no multiple-comparison adjustment.

What I calculate

For a gap δ, the unpaired variance is pA(1 − pA) + pB(1 − pB). For paired items it is q − δ², where q is discordance. Equivalently, subtract twice the covariance from the unpaired variance. I find the smallest integer n with Φ(−z − |δ|√(n/v)) + Φ(|δ|√(n/v) − z) at least the target power, where z is the two-sided normal critical value.

Source, Python comparison fixture, and tests. The browser and Python item counts must agree within one item on the committed input grid. The formula uses population variance implied by your assumptions, not a fitted sample variance.