Alp Cetin

San Francisco

seed-power

I test whether pilot-based seed budgets deliver their promised detection rate in reinforcement learning.

Question

How many independent training seeds does an RL comparison need? If I estimate the gap from a small pilot and plan a larger study around it, does that study detect a difference as often as promised?

Method

I used established power calculations and affine ARS/CEM policy search adapted from Control Clock. I fixed six different-variant comparisons per task on CartPole and Acrobot, changing iteration counts or comparing the optimizers. Each repeat used disjoint eight-seed pilots per arm to estimate the gap and variance, then plan a study targeting 80% power. A policy score averages sixteen evaluation episodes; those episodes are not extra training seeds. The README describes the design and its assumptions.

I committed and pushed the complete pilot plan before confirmation. I then trained 11,440 fresh seeds, separate from 2,688 pilot seeds, and tested the resulting comparisons with Welch’s test. Every executable plan received its exact requested seed count (counts).

Result

The plans targeted 80% power but detected a difference in 67 of 119 executable different-variant comparisons: 56.3%, with a descriptive Wilson 95% interval of 47.3–64.9%. CartPole detected differences in 30/60 comparisons (50.0%); Acrobot in 37/59 (62.7%). This is detection frequency for a fixed set of short-budget comparisons, not power at a known effect.

Detection rates of fresh-seed RL studies planned for 80% powerAll executable plans detected 67 of 119 differences, 56.3%. CartPole detected 30 of 60, 50.0%; Acrobot detected 37 of 59, 62.7%. Bars are Wilson 95% intervals. The dashed line is the 80% planning target.0%20%40%60%80%100%80% targetAll executable plans67/119 (56.3%)CartPole30/60 (50.0%)Acrobot37/59 (62.7%)detection rate; Wilson 95% intervals
Fresh-seed detection rates for executable different-variant plans, compared with the 80% planning target. Intervals are descriptive Wilson 95% intervals. Data: summary.json.

Of 144 attempted different-variant plans, 25 needed more than the predeclared limit of 256 seeds per arm and were not run. I did not cap their counts and still call them 80% plans. Seven detections reversed the pilot’s direction. Separate same-variant controls rejected in 1/17 tests; that small sample cannot establish exact type-I calibration (outcomes).

Limits

These are two cheap control tasks with affine policy search, not neural-network deep RL. The achieved rate is conditional on plans that fit the seed limit. Twelve repeats per comparison leave wide intervals, and pooling different comparisons makes the overall interval descriptive. Fixed iteration budgets are not equal environment-step budgets.

The planner treats noisy pilot gaps and variances as population values. Its equal-variance Gaussian calculation is pooled-t power; using it for Welch’s test is an approximation. Bounded, skewed RL returns need not fit that model, and some different-variant comparisons may have no true mean difference. This result does not show that all missed detections came from pilot noise, or estimate power for RL in general (interpretation and limits).