Alp Cetin

San Francisco

attention-numerics

I measured real FP8 attention in every layer. With FA3 on H100, rotating queries and keys raises Qwen2.5-1.5B’s batch exp-CE ratio to 12.56× versus native GPU BF16. Centering keys first brings it to 1.004×.

Question

Rotating queries and keys can spread outliers before FP8 rounding, but my earlier CPU emulator found heads where it made attention less accurate. Does that harm survive in the published GPU kernels, does it reach the model’s loss, and does the same predictor still identify the harmed heads?

Method

I run actual FlashAttention-3 E4M3 attention on H100 and SageAttention INT8-QK/FP8-PV on L4, using the earlier native BF16 query, key and value captures. I compare each head’s output with ideal causal FP64 attention. The predictor and its zero threshold are unchanged: no kernel-specific refit or tuning. For downstream loss, I replace attention in every layer of two Qwen models with FA3; projections, RoPE, norms and MLP stay native BF16. Each model has 3,072 held-out next-token labels total, from three public-domain books. I report exp(CE_variant − CE_BF16) against matched native GPU BF16 attention: batch teacher forcing, not streaming decoder perplexity. Method and kernel settings.

“Tile” in the result files is the legacy unrotated-condition label. It does not mean FA3 uses the emulator’s per-tile scales: the actual FA3 API uses independent Q/K/V scales over the full sequence for each batch and head. Rotation applies the same Hadamard transform to Q and K, never V. A common vector in K shifts all of a query’s scores equally, which exact softmax cancels. Rotation can spread that vector across channels and make rounding of the token-specific residual worse. Subtracting the mean key preserves exact attention while reducing this numerical problem. Derivation.

Result

The loss increase survives in real FA3. On Qwen2.5-1.5B, the batch exp-CE ratio is 12.56× with Q/K rotation, versus 4.81× unrotated, 1.014× with key centering alone and 1.004× with centering before rotation. On Qwen2.5-0.5B, the corresponding ratios are 2.032×, 1.071×, 1.018× and 1.056×. These are ratios to native GPU BF16, not ratios of cross-entropy itself.

FA3 batch exp-CE ratio versus native GPU BF16 in both Qwen modelsReal FA3 E4M3 attention in every layer on H100, log scale. Qwen2.5-1.5B: unrotated 4.808, rotated 12.560, center K 1.014, rotate plus center K 1.004. Qwen2.5-0.5B: 1.071, 2.032, 1.018 and 1.056 respectively. Each uses 3,072 held-out next-token labels. This is batch teacher forcing, not streaming perplexity.FA3 E4M3 · H100 · all layers · 3,072 labels/model1×2×4×8×16×Qwen2.5-1.5Bunrotated4.808×rotate Q/K12.560×center K1.014×rotate + center K1.004×Qwen2.5-0.5Bunrotated1.071×rotate Q/K2.032×center K1.018×rotate + center K1.056×exp(CE variant − CE BF16), log scale
Actual FA3 E4M3 attention on H100 in every layer of two Qwen models. Batch exp-CE ratio, exp(CE_variant − CE_BF16), versus native GPU BF16, log scale; 3,072 held-out next-token labels per model from three books. Batch teacher forcing, not streaming perplexity. Data: hardware summary.json.

For head-level transfer, I measured all 672 physical query heads in the two Qwens. Rotation harms a head when its mean log1p(relative Frobenius output error) across the three texts exceeds the unrotated condition. The unchanged predictor’s zero threshold gives:

KernelHarmed / 672AUCPrecision / recall
FA3, H100870.95664% / 91%
Sage, L450.9664% / 100%

Sage’s high AUC does not make the threshold useful: it produces 118 false alarms for only five harmed Qwen heads. Across all 918 selected physical heads, the emulator-versus-kernel rotation-effect rank correlation is 0.982 for FA3 but only 0.263 for Sage. The other four models contribute top-ranked and independent uniform samples, with overlaps measured once; the combined sample is enriched, not a population estimate. Head-transfer statistics.

The earlier CPU emulator study covered six models and 2,880 heads. I published the predictor and threshold before loading three evaluation models; AUC was 0.978 on their 1,728 unseen emulated heads. That is an emulator result, not the hardware result above. Sources: earlier study summary and prediction results.

Limits

Six small checkpoints, three books and one 1,024-token context length do not establish task-general behavior. Downstream loss was measured only for FA3 on the two Qwens; Sage has per-head comparisons only. Heads within a model are dependent. Full-batch means and scales may see future tokens, so these results are not generation measurements or latency benchmarks.

API scaling, probability representation, transform rounding and accumulation change together; I cannot isolate one as the cause of kernel differences. The predictor needs ideal attention output, so it is a diagnostic, not a runtime switch. The exact operand bundle is private, not publicly downloadable. The capture pipeline is public and source hashes are published; rebuilding captures with a different library build may change the BF16 operands.