Alp Cetin

San Francisco

control-clock

How many seconds does a fresh process need to learn CartPole or Acrobot on a laptop CPU?

Question

Reinforcement-learning results usually count environment steps. I wanted wall-clock time on an ordinary laptop, measured from the moment the process starts, so that imports, setup and evaluation all count. Does a PPO built for the CPU beat well-tuned libraries, and does PPO beat plain search over linear policies?

Method

Each run starts a fresh Python process and stops the clock when the policy first passes the task’s standard bar: a mean return of at least 475 on CartPole, or −100 on Acrobot, over 100 greedy evaluation episodes whose seeds were never used in training. Every evaluation counts against the clock. The protocol was written down before any tuning.

The methods are my PPO, which steps many environments at once in NumPy; Stable-Baselines3’s PPO with RL Zoo settings; CleanRL’s PPO; and two black-box searches over linear policies, augmented random search (ARS) and the cross-entropy method (CEM). The three methods I wrote ran 20 seeds per task and the two libraries 5, one run at a time on an otherwise idle CPU.

Result

Search over linear policies wins CartPole. ARS passed in a median 1.08 seconds from process start and CEM in 1.15, against 4.15 for my PPO, 4.72 for Stable-Baselines3 and 5.55 for CleanRL. Acrobot reverses the order: Stable-Baselines3 with Zoo settings was fastest at 10.05 seconds, then CleanRL at 15.15 and CEM at 15.34, while my PPO took 25.01 and ARS 28.15. All 140 runs passed within their time limits.

Seconds from process start to a passing policy, per seedMedian seconds by method. CartPole-v1: ARS 1.08 s; CEM 1.15 s; batched PPO 4.15 s; Stable-Baselines3 PPO 4.72 s; CleanRL PPO 5.55 s. Acrobot-v1: Stable-Baselines3 PPO 10.05 s; CleanRL PPO 15.15 s; CEM 15.34 s; batched PPO 25.01 s; ARS 28.15 s. 140 of 140 runs passed.one dot per seed; the bar is the medianpassedCartPole-v1, pass at mean return 475ARS, linear policy20/20CEM, linear policy20/20batched PPO (this repo)20/20SB3 PPO, Zoo settings5/5CleanRL PPO5/5Acrobot-v1, pass at mean return −100SB3 PPO, Zoo settings5/5CleanRL PPO5/5CEM, linear policy20/20batched PPO (this repo)20/20ARS, linear policy20/200.5125102050100seconds from process start, log scale
Seconds from process start to the first passing evaluation, one dot per seed, on a log scale; bars mark medians. The column on the right counts runs that passed. Data: summary.json.

Faster simulation didn’t make my PPO the fastest learner. Its objective, normalization and evaluation schedule all differ from the libraries’, so this comparison can’t say which difference is responsible. CartPole mostly shows that a linear controller is enough for CartPole. One Acrobot run passed before any training, because its random initial policy already averaged −95.6. It counts, as the protocol says, so these are times to qualify, not times spent learning.

Separately, I ran JAX PPO on a cloud NVIDIA L4. With Python startup, CUDA initialization, JIT compilation and every evaluation on the process clock, CartPole passed in a median 26.38 seconds and Acrobot in 97.43 seconds, with 3 of 3 seeds passing each task (GPU summary). Both were slower than the laptop winners. This changes hardware, backend and PPO settings together; it is not an isolated GPU speedup test. The total conservative cloud-cost estimate was about $0.37, including failed setup attempts (cost file).

Limits

One laptop. The libraries ran 5 seeds and the rest 20, so these are descriptive comparisons, not significance tests. The same 100 evaluation seeds decided when to stop and guided development choices, so they are not an untouched test set. LunarLander was planned as a stretch goal and has no final runs.