Verge Lab
I choose preference pairs from aspect scores and abstain when the aspects trade off.
Question
Does refusing aspect tradeoffs select better human preference labels than choosing pairs with a large overall-score gap? I keep a pair only when one response is at least as good on every chosen aspect and better on at least one, rather than averaging away disagreement.
Method
I joined HelpSteer2’s human aspect ratings to its separately collected pairwise preferences by exact prompt and response hashes. My primary Pareto rule compares correctness and coherence. I hold out overall helpfulness for the baseline, which selects the largest absolute helpfulness gaps and predicts their direction. I exclude helpfulness ties from both selectors, match their selected pair counts, and evaluate against the separate human preferences. Equal-gap tie-breaking does not use preference labels.
I take the normalized ratings at face value, with no uncertainty penalty. Complexity and verbosity describe style, so I do not maximize them in the primary rule. Kept pairs can be exported for DPO, but I did not train a model.
Result
I found no clear advantage on HelpSteer2 validation. Both selectors kept 310 pairs. Pareto agreed with 281/286 strict human preferences (98.25%); the matched overall-score-gap baseline agreed with 277/282 (98.23%). Both contradicted 5 human preferences.
Equal yield does not mean identical pairs or equal numbers of strict human labels. Pareto’s selected pairs include 24 human ties, leaving 286 strict preferences; the baseline’s include 28, leaving 282. I do not count ties as agreement. The baseline also varies with label-blind tie-breaking, so the tiny headline difference does not establish superiority.
Earlier, I tested the first 200 TruthfulQA prompts in UltraFeedback, using GPT-4 aspect ratings. I kept 820 of 1,200 pairs. Of the kept pairs with a strict separate overall ranking, 132/706 (18.7%) went the other way. This measures inconsistency between LLM ratings, not human preference errors.
Limits
I tested a small HelpSteer2 validation sample. Its human ratings are rounded, filtered annotations, not objective truth; separately collected preferences do not guarantee independent annotators. The earlier UltraFeedback sample is a deterministic prefix, not a random sample, and uses unverified GPT-4 ratings rather than human labels. I measured agreement with recorded preferences, not downstream training benefit. The browser review tool uses authored illustrative examples, not these benchmark rows.
Links
Code and results GitHub · pinned revision
Interactive review tool illustrative examples, in the browser
HelpSteer2 · HelpSteer2-Preference human ratings and preferences
UltraFeedback Cui et al., 2023