1. Plan a comparison
Accuracies are binary correct/incorrect rates. For unpaired studies, the budget is per model; for paired studies, both models answer the same items.
Items needed
—
per model
Choose the comparison, then calculate.
An approximation, not a guarantee. I assume independent, identically distributed items and a fixed variance. This is not exact McNemar power, a clustered-item budget, a finite-pool calculation, or a multiple-comparison correction.
For very small budgets or rare discordance, the Gaussian approximation can be poor. Keep the meaningful gap fixed; do not use a noisy pilot gap as a promise.
2. Did pilot plans deliver 80% power?
I checked plans from small pilots against disjoint heldout items in five historical benchmarks. This panel is a fixed result from committed data, not a prediction for the inputs above.
Loading committed results…
Each benchmark has 40 score-independent model pairs × 5 item splits. Each evaluable plan has 1,000 exact two-sided McNemar trials from the empirical heldout distribution, sampled with replacement. An estimated detection rate is not the true power; Monte Carlo bounds show simulation uncertainty only.
These are selected 395-model, 2024 Open LLM Leaderboard data, not current frontier models. Resampling existing items does not establish power on newly collected questions. The 80% target uses alpha 0.05, with no multiple-comparison adjustment.
Summary CSV · Counts derived from individual plans · Committed individual plans · Experiment settings
What I calculate
For a gap δ, the unpaired variance is pA(1 − pA) + pB(1 − pB). For paired items it is q − δ², where q is discordance. Equivalently, subtract twice the covariance from the unpaired variance. I find the smallest integer n with Φ(−z − |δ|√(n/v)) + Φ(|δ|√(n/v) − z) at least the target power, where z is the two-sided normal critical value.
Source, Python comparison fixture, and tests. The browser and Python item counts must agree within one item on the committed input grid. The formula uses population variance implied by your assumptions, not a fitted sample variance.