control-clock
How many seconds does a fresh process need to learn CartPole or Acrobot on a laptop CPU?
Question
Reinforcement-learning results usually count environment steps. I wanted wall-clock time on an ordinary laptop, measured from the moment the process starts, so that imports, setup and evaluation all count. Does a PPO built for the CPU beat well-tuned libraries, and does PPO beat plain search over linear policies?
Method
Each run starts a fresh Python process and stops the clock when the policy first passes the task’s standard bar: a mean return of at least 475 on CartPole, or −100 on Acrobot, over 100 greedy evaluation episodes whose seeds were never used in training. Every evaluation counts against the clock. The protocol was written down before any tuning.
The methods are my PPO, which steps many environments at once in NumPy; Stable-Baselines3’s PPO with RL Zoo settings; CleanRL’s PPO; and two black-box searches over linear policies, augmented random search (ARS) and the cross-entropy method (CEM). The three methods I wrote ran 20 seeds per task and the two libraries 5, one run at a time on an otherwise idle CPU.
Result
Search over linear policies wins CartPole. ARS passed in a median 1.08 seconds from process start and CEM in 1.15, against 4.15 for my PPO, 4.72 for Stable-Baselines3 and 5.55 for CleanRL. Acrobot reverses the order: Stable-Baselines3 with Zoo settings was fastest at 10.05 seconds, then CleanRL at 15.15 and CEM at 15.34, while my PPO took 25.01 and ARS 28.15. All 140 runs passed within their time limits.
Faster simulation didn’t make my PPO the fastest learner. Its objective, normalization and evaluation schedule all differ from the libraries’, so this comparison can’t say which difference is responsible. CartPole mostly shows that a linear controller is enough for CartPole. One Acrobot run passed before any training, because its random initial policy already averaged −95.6. It counts, as the protocol says, so these are times to qualify, not times spent learning.
Separately, I ran JAX PPO on a cloud NVIDIA L4. With Python startup, CUDA initialization, JIT compilation and every evaluation on the process clock, CartPole passed in a median 26.38 seconds and Acrobot in 97.43 seconds, with 3 of 3 seeds passing each task (GPU summary). Both were slower than the laptop winners. This changes hardware, backend and PPO settings together; it is not an isolated GPU speedup test. The total conservative cloud-cost estimate was about $0.37, including failed setup attempts (cost file).
Limits
One laptop. The libraries ran 5 seeds and the rest 20, so these are descriptive comparisons, not significance tests. The same 100 evaluation seeds decided when to stop and guided development choices, so they are not an untouched test set. LunarLander was planned as a stretch goal and has no final runs.
Links
Code, data and results GitHub
Simple random search provides a competitive approach to reinforcement learning Mania, Guy and Recht, 2018
RL Baselines3 Zoo Raffin, 2020
CleanRL Huang et al., 2022