BranchPilot
When has a language model sampled enough answers?
Question
Self-consistency asks a model the same question several times and keeps the majority answer. Each extra sample costs time and money, and most problems don’t need all of them. BranchPilot decides after every answer whether to ask again or stop. I wanted to know whether a learned stopping rule beats simple ones.
Method
After each sample, a strategy sees only the observed prefix, such as vote counts and whether the latest answers agree, and returns stop or continue. I compared fixed sample counts, vote-share thresholds, stopping after two or three matching answers in a row, and a learned rule. A small network trained on 1,600 logged GSM8K problems predicts whether another sample is worth a specified penalty, at six penalties per sample. These penalties are not dollar prices.
The answers came from Qwen2.5-1.5B-Instruct on an NVIDIA L4, eight samples for each of 1,600 training, 400 validation and all 1,319 test problems. Every rule’s settings were chosen on validation before the test run, and the success criterion was written down before it.
For the second benchmark I collected 4,000 new responses from the same model and fixed an internal MATH-500 split of 200 training, 100 validation and 200 held-out problems. I retrained the controller and separately transferred the old one unchanged. Symbolic voting used earlier responses only; gold correctness labels stayed separate. The protocol preceded sampling, and the simple comparators were chosen on validation before holdout evaluation.
Result
On GSM8K, the learned rule failed the success criterion. Against the simple rule chosen on validation at each penalty, it did measurably worse at five of the six penalties and tied at the sixth.
It did beat fixed sample counts at small budgets: 72.5% accuracy with 1.5 samples per problem, against 71.9% for always drawing two. But stopping after two matching answers reached 78.8% with 3.2 samples, and the learned rule never caught up.
The negative result extends to the internal 200-problem MATH-500 holdout. At a sample penalty of 0.05, the retrained controller reached 58.5% accuracy at 3.35 samples per problem. Stopping after two matching answers reached 62.5% at 4.33 samples; always drawing eight also reached 62.5%. These accuracy comparisons are descriptive, not the fixed success test.
The fixed objective was accuracy − penalty × (samples − 1). Validation selected always drawing one sample as the comparator at all three primary penalties. The retrained controller’s paired utility differences were −0.0825, −0.1048 and −0.0955, with all three 95% intervals below zero. It failed the unchanged success rule on this second benchmark too.
At the 0.05 penalty and exactly matched expected sample budgets, neither paired accuracy interval against simple mixtures established an advantage for the retrained controller. Those mixture weights use holdout sample counts, not accuracy, and their uncertainty is not included in the intervals.
Limits
One small model and one generation seed per study, with at most eight samples per problem. The MATH result is an internal 200-problem holdout, not a full official MATH-500 score. Symbolic parsing can fail; outputs without answer labels retain the original unknown/tie handling. This does not show that learned stopping is generally impossible.
Response-bank replay is not live-serving latency. Sample counts are not token, energy or dollar budgets. Intervals are pointwise and conditional on the bank and fitted policy, not variation across training or generation seeds. Training targets use gold and logged future answers; they are exact for those trajectories, not an optimal stopping solution.
Links
Code and results GitHub
New response bank and generation settings Pinned source
Adaptive-Consistency Aggarwal et al., 2023
Early-Stopping Self-Consistency Li et al., 2024