Alp Cetin

San Francisco

Verge Lab

I choose preference pairs from aspect scores and abstain when the aspects trade off.

Question

Does refusing aspect tradeoffs select better human preference labels than choosing pairs with a large overall-score gap? I keep a pair only when one response is at least as good on every chosen aspect and better on at least one, rather than averaging away disagreement.

Method

I joined HelpSteer2’s human aspect ratings to its separately collected pairwise preferences by exact prompt and response hashes. My primary Pareto rule compares correctness and coherence. I hold out overall helpfulness for the baseline, which selects the largest absolute helpfulness gaps and predicts their direction. I exclude helpfulness ties from both selectors, match their selected pair counts, and evaluate against the separate human preferences. Equal-gap tie-breaking does not use preference labels.

I take the normalized ratings at face value, with no uncertainty penalty. Complexity and verbosity describe style, so I do not maximize them in the primary rule. Kept pairs can be exported for DPO, but I did not train a model.

Result

I found no clear advantage on HelpSteer2 validation. Both selectors kept 310 pairs. Pareto agreed with 281/286 strict human preferences (98.25%); the matched overall-score-gap baseline agreed with 277/282 (98.23%). Both contradicted 5 human preferences.

Agreement with strict human preferences at matched yieldAt 310 selected pairs each, Pareto selection agreed with 281 of 286 strict preferences, 98.25%, and the overall-score-gap baseline with 277 of 282, 98.23%. Wilson 95% intervals overlap. Human ties are excluded from agreement.95%96%97%98%99%100%Pareto selection281/286 (98.25%)Overall-score gap277/282 (98.23%)strict agreement; Wilson 95% intervals
Human preference agreement on HelpSteer2 validation at matched yield. Human ties stay in the selected count, but not the strict agreement denominator. Data: human-preferences/summary.json.

Equal yield does not mean identical pairs or equal numbers of strict human labels. Pareto’s selected pairs include 24 human ties, leaving 286 strict preferences; the baseline’s include 28, leaving 282. I do not count ties as agreement. The baseline also varies with label-blind tie-breaking, so the tiny headline difference does not establish superiority.

Earlier, I tested the first 200 TruthfulQA prompts in UltraFeedback, using GPT-4 aspect ratings. I kept 820 of 1,200 pairs. Of the kept pairs with a strict separate overall ranking, 132/706 (18.7%) went the other way. This measures inconsistency between LLM ratings, not human preference errors.

Outcomes for 1,200 response pairs from UltraFeedbackWith point scores, 820 pairs are defended and 380 abstained. Of the defended pairs, 574 agree with the dataset's overall score, 132 disagree and 114 have tied overall scores. Assuming judge confidence 0.9, 343 are defended.point scores, 1,200 pairs574 defended, agree with overall score132114380 abstaineddisagreeoverall tieassuming judge confidence 0.9, uncertainty penalty 0.05343 defended857 abstained
Earlier UltraFeedback outcomes against the separate GPT-4 overall score. The lower bar is a modeling sensitivity: assumed confidence 0.9 with uncertainty penalty scale 0.05, not measured judge reliability. Data: public-preferences/summary.json.

Limits

I tested a small HelpSteer2 validation sample. Its human ratings are rounded, filtered annotations, not objective truth; separately collected preferences do not guarantee independent annotators. The earlier UltraFeedback sample is a deterministic prefix, not a random sample, and uses unverified GPT-4 ratings rather than human labels. I measured agreement with recorded preferences, not downstream training benefit. The browser review tool uses authored illustrative examples, not these benchmark rows.