attention-numerics
I measured real FP8 attention in every layer. With FA3 on H100, rotating queries and keys raises Qwen2.5-1.5B’s batch exp-CE ratio to 12.56× versus native GPU BF16. Centering keys first brings it to 1.004×.
Question
Rotating queries and keys can spread outliers before FP8 rounding, but my earlier CPU emulator found heads where it made attention less accurate. Does that harm survive in the published GPU kernels, does it reach the model’s loss, and does the same predictor still identify the harmed heads?
Method
I run actual FlashAttention-3 E4M3 attention on H100 and SageAttention INT8-QK/FP8-PV on L4, using the earlier native BF16 query, key and value captures. I compare each head’s output with ideal causal FP64 attention. The predictor and its zero threshold are unchanged: no kernel-specific refit or tuning. For downstream loss, I replace attention in every layer of two Qwen models with FA3; projections, RoPE, norms and MLP stay native BF16. Each model has 3,072 held-out next-token labels total, from three public-domain books. I report exp(CE_variant − CE_BF16) against matched native GPU BF16 attention: batch teacher forcing, not streaming decoder perplexity. Method and kernel settings.
“Tile” in the result files is the legacy unrotated-condition label. It does not mean FA3 uses the emulator’s per-tile scales: the actual FA3 API uses independent Q/K/V scales over the full sequence for each batch and head. Rotation applies the same Hadamard transform to Q and K, never V. A common vector in K shifts all of a query’s scores equally, which exact softmax cancels. Rotation can spread that vector across channels and make rounding of the token-specific residual worse. Subtracting the mean key preserves exact attention while reducing this numerical problem. Derivation.
Result
The loss increase survives in real FA3. On Qwen2.5-1.5B, the batch exp-CE ratio is 12.56× with Q/K rotation, versus 4.81× unrotated, 1.014× with key centering alone and 1.004× with centering before rotation. On Qwen2.5-0.5B, the corresponding ratios are 2.032×, 1.071×, 1.018× and 1.056×. These are ratios to native GPU BF16, not ratios of cross-entropy itself.
For head-level transfer, I measured all 672 physical query heads in the two Qwens. Rotation harms a head when its mean log1p(relative Frobenius output error) across the three texts exceeds the unrotated condition. The unchanged predictor’s zero threshold gives:
| Kernel | Harmed / 672 | AUC | Precision / recall |
|---|---|---|---|
| FA3, H100 | 87 | 0.956 | 64% / 91% |
| Sage, L4 | 5 | 0.966 | 4% / 100% |
Sage’s high AUC does not make the threshold useful: it produces 118 false alarms for only five harmed Qwen heads. Across all 918 selected physical heads, the emulator-versus-kernel rotation-effect rank correlation is 0.982 for FA3 but only 0.263 for Sage. The other four models contribute top-ranked and independent uniform samples, with overlaps measured once; the combined sample is enriched, not a population estimate. Head-transfer statistics.
The earlier CPU emulator study covered six models and 2,880 heads. I published the predictor and threshold before loading three evaluation models; AUC was 0.978 on their 1,728 unseen emulated heads. That is an emulator result, not the hardware result above. Sources: earlier study summary and prediction results.
Limits
Six small checkpoints, three books and one 1,024-token context length do not establish task-general behavior. Downstream loss was measured only for FA3 on the two Qwens; Sage has per-head comparisons only. Heads within a model are dependent. Full-batch means and scales may see future tokens, so these results are not generation measurements or latency benchmarks.
API scaling, probability representation, transform rounding and accumulation change together; I cannot isolate one as the cause of kernel differences. The predictor needs ideal attention output, so it is a diagnostic, not a runtime switch. The exact operand bundle is private, not publicly downloadable. The capture pipeline is public and source hashes are published; rebuilding captures with a different library build may change the BF16 operands.
Links
Code, data and results GitHub
Real-kernel method and reproduction Pinned source
Explore the earlier CPU emulator and rounding mechanism Interactive
FlashAttention-3 Shah et al., 2024
SageAttention Zhang et al., 2025
SageAttention2 Zhang et al., 2025
QuaRot Ashkboos et al., 2024