Alp Cetin

San Francisco

AlignmentTax

Does instruction tuning trade calibration for truthfulness?

Question

A model can answer more questions correctly while becoming too confident about the answers it gets wrong. I asked whether instruction tuning consistently makes that trade: better truthfulness, worse confidence calibration.

Method

I compared seven matched base/instruct pairs across four families: Qwen2.5 at 0.5B, 1.5B and 7B; OLMo-2 at 1B and 7B; SmolLM2-1.7B; and Mistral-7B-v0.3. Each checkpoint answered the same 790 questions from a binary TruthfulQA derivative. I used shared plain prompts and normalized the likelihoods of the A/B answer labels, not generated explanations or verbal confidence.

I measured accuracy and expected calibration error (ECE), using ten equal-frequency bins. The 95% intervals use 10,000 paired question resamples, recomputing ECE each time. The analysis plan was committed before scoring; model and tokenizer commits are pinned.

Result

Calibration usually worsens, but not consistently in exchange for better truthfulness. Four pairs gain accuracy and worsen ECE; two lose accuracy and worsen ECE, with both intervals excluding zero. Qwen2.5-0.5B loses accuracy, with an uncertain ECE change.

Instruction tuning changes binary accuracy and calibration with shared promptsIn seven base/instruct pairs, four gain accuracy and increase expected calibration error; two lose accuracy and increase error. Qwen2.5-0.5B loses accuracy, with an uncertain calibration change. SmolLM2-1.7B loses 15.8 accuracy points and gains 32.5 error points. Horizontal and vertical bars are separate 95% question-bootstrap intervals, not a joint confidence region.shared prompts; instruct − base; 95% paired intervalsECE change (points); higher is worse010203040-20-10010accuracy change (percentage points)11 Qwen2.5-0.5B22 Qwen2.5-1.5B33 Qwen2.5-7B44 OLMo-2 1B (0425)55 OLMo-2 7B (1124)66 SmolLM2-1.7B77 Mistral-7B-v0.3
Instruct minus base under shared prompts. Right means higher binary accuracy; up means worse ECE. Bars are separate 95% paired-bootstrap intervals for each metric, not a joint confidence region. Data: summary.json, full metric table.

SmolLM2-1.7B loses 15.8 percentage points of accuracy and adds 32.5 points of ECE. That is not a trade for better answers: both measures get worse. The metric also matters. Qwen2.5-7B’s Brier score improves while its ECE and negative log-likelihood worsen, so I report all the metrics rather than treating them as interchangeable.

I also checked native chat templates. That comparison changes both the checkpoint and its prompt; it cannot isolate a weights-only effect. The figure above uses only the shared-prompt comparison.

Limits

This is one binary TruthfulQA derivative, not official MC1/MC2 or free-form truthfulness. Confidence is restricted to the A/B labels, not all possible answers. Base/instruct checkpoints differ in training stages and data; this is not a causal effect of reinforcement learning.

These selected checkpoints are not a model-family population. The intervals resample questions, not training seeds, and are pointwise rather than adjusted for multiple comparisons. ECE depends on binning. I used scalar bf16 inference after a batching pilot failed its numerical tolerance; no precision sweep was run.