Faultline
Can a reinforcement-learning agent learn to run a cheap test before an expensive repair?
Question
In real debugging, different faults often look the same until you run a test. Faultline is a small factory simulator built around that: two hidden faults produce identical symptoms but need opposite repairs, and a cheap inspection tells them apart. I wanted to know whether training agents only on these ambiguous cases teaches them to diagnose better than a 50/50 mix of ambiguous and obvious cases, or a curriculum that adapts to the agent’s failures.
Method
The agent is a graph encoder with a recurrent core, trained with PPO. Its reward is production minus operating costs, with no bonus for gathering information, so an inspection has to pay for itself. I compare a 50/50 mix of ambiguous and revealed tasks, a difficulty-adaptive mix, and ambiguous tasks only.
For the independent confirmation, I fixed 375 matched training seeds, 500–874, before training: 1,125 policies, each targeting 30,000 decisions. Every run evaluates the same 128 validation base pairs. Diagnostic success requires advancing, obtaining informative inspection evidence, then making the correct repair; guessing alone does not count. The 95% intervals use 10,000 bootstrap resamples over training seeds, not episodes. Poor learning runs remain included.
Result
Ambiguous-only training beats random sampling, but loses to the difficulty curriculum. Mean diagnostic success is 73.8% for Random, 89.9% for Difficulty, and 83.9% for Epistemic (all ambiguous).
The paired Epistemic−Random difference is +10.1 percentage points, with a 95% interval of [+4.7, +15.5]. Epistemic−Difficulty is −5.9 points [−10.3, −1.7]. The hoped-for advantage over both is ruled out. These are individual intervals, not simultaneous intervals for a full curriculum ranking. Neither establishes a minimum five-point effect or practical equivalence within ±5 points.
The earlier 32-seed replication was inconclusive: both primary paired intervals crossed zero. I do not pool it with this confirmation.
Limits
One synthetic task family and a small policy, evaluated on the validation split only; the held-out test split remains untouched. The observations do not include cue reliability or reward coefficients, so this does not establish when agents adapt testing to reliability or cost.
The plan targeted interval half-widths below five points. I missed that target against Random (5.43 points) and did not extend the cohort after seeing outcomes. The confirmation used one PyTorch thread; earlier cohorts used more, so I do not assume identical arithmetic or trajectories. Earlier evidence-swapping tests concern older policies, not these new checkpoints.
Links
Independent confirmation plan Pinned source