attention-numerics / Alp Cetin

A rounding problem, not a rotation problem

When rotation hurts attention

Rotation preserves dot products. Rounding does not. I measured what happens when the two meet in low-precision attention.

250 / 2,880heads made less accurate
0.978 AUCprediction on three unseen models
1.91×Qwen2.5-0.5B perplexity with rotation

These are CPU FP8 emulation results across six small open models, not measurements of a FlashAttention-3 GPU kernel. Below, I separate an illustration of the mechanism from the measured heads.

01 / Synthetic illustration

A shared key component should disappear

I give four keys the same first coordinate. It adds the same constant to every attention score, so exact softmax cancels it. Move the slider: the exact weights stay fixed. After rotation spreads that common component across channels, coarse FP8 rounding can erase the small differences that softmax needs.

Attention weights

━ Exact━ FP8 without rotation━ FP8 with rotation

Weight error

Total variation: the fraction of attention mass moved.

Without rotation
—
With rotation
—

See the vectors and the arithmetic

One query, four keys, eight channels. Q and K each use one maximum-based FP8 scale. I apply a fixed random-sign Hadamard to both before rounding; the transform is orthogonal. E4M3FN uses round-to-nearest-even and saturation at ±448. Dot products and softmax use JavaScript numbers; Q/K scaling and conversion use float32.

This panel stops at attention weights: it does not round probabilities or values, model causal tiles, or reproduce the measured output errors below. Mean subtraction is before rotation and leaves exact attention unchanged.

Key vectors before rounding and dequantized FP8 vectors
KeyOriginalFP8, no rotationFP8, rotated coordinates

02 / Measured heads

Could I predict which heads rotation would hurt?

I fixed a rounding-noise predictor before loading three unseen models. Each point is one physical query head, averaged over three 1,024-token English texts. Both axes compare rotated against unrotated relative attention-output error: positive means rotation hurts; negative means it helps. The zero threshold was fixed, not tuned here.

Loading the committed measurements…

Predicted versus observed effect of rotationOrange points have positive observed harm; blue points do not. Click a point to inspect it, or use the paginated table below.

● Rotation hurts● Rotation helps or ties

What do these scores mean?

For each text I take log(1 + relative Frobenius output error), average those values, then subtract the unrotated average from the rotated one. I do the same with the predictor's errors. These are differences of log-transformed errors, not error ratios or perplexity changes. The inspector converts the averaged log error back with exp(mean) − 1.

The predictor needs ideal attention outputs. It is a diagnostic, not a free runtime switch. Heads share layers and models: I do not treat 2,880 heads as independent experiments. AUC measures ranking, not calibration. OLMo has no harmed heads, so its own AUC is undefined.

Inspect with the keyboard

Rows are ordered by observed harm, largest first. Filters also apply to this table.

Measured heads, 20 per page
ModelLayer / headPredicted scoreObserved scoreInspect

Rotation is not always the right fix

FlashAttention-3 rotates Q and K to spread outliers before FP8 rounding. My results show a failure mode: a large shared vector can look like an outlier, yet its exact score contribution cancels. Spreading it can make token-specific differences harder to represent. Subtracting the mean key is a different intervention; in Qwen2.5-0.5B it brought the rotated perplexity ratio from 1.91× to 1.06×.

I tested six small models and three texts with CPU emulation. These results do not establish the size of this effect in an actual FP8 GPU implementation or in larger models. Read the study, limitations and reproduction commands.

Built on FlashAttention-3 and SageAttention. Data export · FP8 parity tests.