A rounding problem, not a rotation problem
When rotation hurts attention
Rotation preserves dot products. Rounding does not. I measured what happens when the two meet in low-precision attention.
These are CPU FP8 emulation results across six small open models, not measurements of a FlashAttention-3 GPU kernel. Below, I separate an illustration of the mechanism from the measured heads.
01 / Synthetic illustration
A shared key component should disappear
I give four keys the same first coordinate. It adds the same constant to every attention score, so exact softmax cancels it. Move the slider: the exact weights stay fixed. After rotation spreads that common component across channels, coarse FP8 rounding can erase the small differences that softmax needs.
Attention weights
━ Exact━ FP8 without rotation━ FP8 with rotation
Weight error
Total variation: the fraction of attention mass moved.
- Without rotation
- —
- With rotation
- —
See the vectors and the arithmetic
One query, four keys, eight channels. Q and K each use one maximum-based FP8 scale. I apply a fixed random-sign Hadamard to both before rounding; the transform is orthogonal. E4M3FN uses round-to-nearest-even and saturation at ±448. Dot products and softmax use JavaScript numbers; Q/K scaling and conversion use float32.
This panel stops at attention weights: it does not round probabilities or values, model causal tiles, or reproduce the measured output errors below. Mean subtraction is before rotation and leaves exact attention unchanged.
| Key | Original | FP8, no rotation | FP8, rotated coordinates |
|---|
02 / Measured heads
Could I predict which heads rotation would hurt?
I fixed a rounding-noise predictor before loading three unseen models. Each point is one physical query head, averaged over three 1,024-token English texts. Both axes compare rotated against unrotated relative attention-output error: positive means rotation hurts; negative means it helps. The zero threshold was fixed, not tuned here.
Loading the committed measurements…
● Rotation hurts● Rotation helps or ties
What do these scores mean?
For each text I take log(1 + relative Frobenius output error), average those values, then subtract the unrotated average from the rotated one. I do the same with the predictor's errors. These are differences of log-transformed errors, not error ratios or perplexity changes. The inspector converts the averaged log error back with exp(mean) − 1.
The predictor needs ideal attention outputs. It is a diagnostic, not a free runtime switch. Heads share layers and models: I do not treat 2,880 heads as independent experiments. AUC measures ranking, not calibration. OLMo has no harmed heads, so its own AUC is undefined.
Inspect with the keyboard
Rows are ordered by observed harm, largest first. Filters also apply to this table.
| Model | Layer / head | Predicted score | Observed score | Inspect |
|---|
Rotation is not always the right fix
FlashAttention-3 rotates Q and K to spread outliers before FP8 rounding. My results show a failure mode: a large shared vector can look like an outlier, yet its exact score contribution cancels. Spreading it can make token-specific differences harder to represent. Subtracting the mean key is a different intervention; in Qwen2.5-0.5B it brought the rotated perplexity ratio from 1.91× to 1.06×.
I tested six small models and three texts with CPU emulation. These results do not establish the size of this effect in an actual FP8 GPU implementation or in larger models. Read the study, limitations and reproduction commands.
Built on FlashAttention-3 and SageAttention. Data export · FP8 parity tests.