AlignmentTax
Does instruction tuning trade calibration for truthfulness?
Question
A model can answer more questions correctly while becoming too confident about the answers it gets wrong. I asked whether instruction tuning consistently makes that trade: better truthfulness, worse confidence calibration.
Method
I compared seven matched base/instruct pairs across four families: Qwen2.5 at 0.5B, 1.5B and 7B; OLMo-2 at 1B and 7B; SmolLM2-1.7B; and Mistral-7B-v0.3. Each checkpoint answered the same 790 questions from a binary TruthfulQA derivative. I used shared plain prompts and normalized the likelihoods of the A/B answer labels, not generated explanations or verbal confidence.
I measured accuracy and expected calibration error (ECE), using ten equal-frequency bins. The 95% intervals use 10,000 paired question resamples, recomputing ECE each time. The analysis plan was committed before scoring; model and tokenizer commits are pinned.
Result
Calibration usually worsens, but not consistently in exchange for better truthfulness. Four pairs gain accuracy and worsen ECE; two lose accuracy and worsen ECE, with both intervals excluding zero. Qwen2.5-0.5B loses accuracy, with an uncertain ECE change.
SmolLM2-1.7B loses 15.8 percentage points of accuracy and adds 32.5 points of ECE. That is not a trade for better answers: both measures get worse. The metric also matters. Qwen2.5-7B’s Brier score improves while its ECE and negative log-likelihood worsen, so I report all the metrics rather than treating them as interchangeable.
I also checked native chat templates. That comparison changes both the checkpoint and its prompt; it cannot isolate a weights-only effect. The figure above uses only the shared-prompt comparison.
Limits
This is one binary TruthfulQA derivative, not official MC1/MC2 or free-form truthfulness. Confidence is restricted to the A/B labels, not all possible answers. Base/instruct checkpoints differ in training stages and data; this is not a causal effect of reinforcement learning.
These selected checkpoints are not a model-family population. The intervals resample questions, not training seeds, and are pointwise rather than adjusted for multiple comparisons. ECE depends on binning. I used scalar bf16 inference after a batching pilot failed its numerical tolerance; no precision sweep was run.
Links
Code, scores and results GitHub
Scoring and comparison protocol Pinned source
TruthfulQA Lin, Hilton and Evans, 2021
On calibration of modern neural networks Guo et al., 2017
Post-training calibration Huang, Lu and Zeng, 2025