Alp Cetin

San Francisco

HeliosTune

The action set limits H100 tuning: torch.matmul has lower stored median latency on all 96 workloads, even against the fastest of 36 curated Triton actions.

Question

I originally asked whether tuning data from L4, A10 and T4 GPUs could reduce the probes needed on an H100. This audit asks a more basic question: could better search over the existing actions erase the measured gap to torch.matmul?

Method

I audited old stored timings for 96 workloads covering 84 unique matrix shapes from four model families, with 36 curated Triton actions. The historical collector timed warmed, allocating FP16 calls; this is a descriptive reanalysis, not a new timing collection. I select each workload’s reference configuration on bank 1 and score it and torch.matmul on bank 2.

The audit ratio is reference_latency / torch_latency; above one favors torch. I take the geometric mean over workloads. The reference is the best curated action selected on bank 1, not a hardware ceiling. I also inspect the fastest action on bank 2 itself as an optimistic diagnostic.

The original policy uses nearby-shape retrieval and Bayesian linear Thompson sampling. Parhelion starts with a retrieved recommendation, then adapts. Test workloads and their model families were held out of the retrieval archive. Its probe budgets replay exhaustive measurements; they do not measure reduced physical acquisition cost.

Result

On the H100, torch.matmul has lower stored median latency than the independently selected reference on all 96 workloads. The geometric-mean reference_latency / torch_latency ratio is 1.610016×. Even choosing each workload’s fastest of the 36 configurations on the scoring bank cannot beat torch on any workload. That is an optimistic, in-sample diagnostic, not an independently selected policy. Within this fixed timing matrix, a better search policy cannot find a faster-than-torch action.

Stored H100 reference-to-torch latency ratios by matrix row countIn each of six row-count groups, torch.matmul has lower stored median latency than the bank-1-selected Triton reference on all 16 workloads. Geometric-mean reference-to-torch ratios range from 1.404 to 1.680. These are old measurements reanalyzed, without uncertainty estimates or new timings.reference latency / torch latency; higher favors torchstored H100 medians; 16 torch wins / 16 workloads per group1.0×1.2×1.4×1.6×1.8×1.645×11.652×71.678×311.680×961.619×2571.404×1024M in A[M,K] × B[K,N]; geometric-mean ratio
H100 geometric-mean reference_latency / torch_latency by M in A[M,K] × B[K,N], for M = 1, 7, 31, 96, 257 and 1024. Each group has 16 workloads, all with lower stored torch medians. Reference selection uses bank 1; scoring uses bank 2. Data: action-set-audit.json.

The separate GPU overview counts lower torch medians on 31/96 L4, 31/96 A10, 94/96 T4 and 96/96 H100 workloads. L4 and A10 are mixed, not uniformly dominated. These are within-GPU summaries from separate acquisition stages, not a pooled four-GPU policy result.

The historical transfer comparison is smaller: cold-start Thompson sampling scored 0.958434 against Parhelion’s 0.950259, averaged over budgets 1–8. Here score means reference_latency / selected_latency, so higher is better. The original aggregation averages model-family geometric means over policy seeds, then averages the eight budgets; it differs from the audit’s single workload geometric mean.

H100 tuning quality against the number of configurations timedScore is bank-1-selected curated-reference latency divided by selected latency; higher is better, not a fraction of a hardware ceiling. Cold-start Thompson sampling and Parhelion both reach 0.997 by eight probes; reuse and retrieval baselines stay lower. torch.matmul scores 1.61 on the same scale, off the top of the chart.cold-start Thompson samplingParhelion (retrieval first)reuse the nearest measured shaperetrieval onlytorch.matmul: 1.61 ↑curated-reference latency / selected latency0.800.850.900.951.0012345678configurations timed on the H100
Historical H100 reference-relative score by replayed probe budget, averaged over held-out workloads and 30 policy seeds. The curated reference is selected on bank 1 and scored on bank 2; it is not a ceiling. Data: parhelion-h100-final.json.

Both selected transfer strengths were zero. Parhelion still uses a retrieval anchor and source features: this does not show retrieval is useless. Retrieval gave it a better first probe, but cold-start search overtook it on the second.

Limits

Torch’s within-run spreads were not stored, so the positive counts are not significance tests or proof of a robust advantage on every workload. One near-tie saves only 0.096 µs. Policy-seed intervals are not hardware-repeat uncertainty, and repeated shapes across models are not independent devices.

Old warmed medians omit compilation, policy computation and serving overhead. Timed torch FP16 accumulation settings were not fully recorded. Correctness used a separate FP32-input torch reference with TF32 disabled, converted to FP16; that check does not prove matched accumulation in the timed comparison.

This fixed corpus does not identify a causal hardware bottleneck or show that a broader Triton action set would lose. The separate H200 engineering run used a different protocol and is not pooled here.