hardware-counters
noncircular.txt
Non-circular cross-machine projection: ARCHER2 -> Cirrus
==============================================================================
ARCHER2 merged configurations : 191
Cirrus merged configurations : 191
matched on app/size/ncore/nthread/nrank : 166
applications on both machines : 6 (comd, hpcg, hpl, lulesh, minife, stream)
gromacs and openfoam are ARCHER2-only and are excluded.
features: 27, all derived from ARCHER2 plus static specs
NO Cirrus counter and NO Cirrus runtime enters any feature.
app n med t_A2 med t_CIR speedup sd log10
comd 25 2.091 0.525 3.647 0.093
hpcg 44 39.645 33.356 1.182 0.075
hpl 24 1.633 0.451 4.131 0.145
lulesh 13 5.342 0.928 4.322 0.132
minife 36 0.624 0.375 2.736 0.240
stream 24 3.313 1.462 2.243 0.027
ALL 166 3.378 1.203 2.289 0.244
clock ratio f_cir/f_a2 = 1.649, so a pure clock argument
predicts a uniform 1.649x speedup.
--- leave-one-application-out, error = max(pred/act, act/pred) ---
predictor kind median p90 macro
t_cir = t_a2 (no adjustment) reference 2.289 4.469 3.043
t_cir = t_a2 * f_a2/f_cir (clock ratio) reference 1.608 2.722 1.990
median training Cirrus runtime reference 5.086 71.646 14.571
t_a2 * constant speedup (fitted) 1 parameter 1.714 3.079 1.831
RF on A2 counters -> speedup learned 1.637 3.445 1.673
Ridge on A2 counters -> speedup learned 2.508 12.953 4.706
RF on A2 counters -> log t_cir learned 1.735 6.328 2.255
A2 cycles / f_cir, constant eta learned 1.841 3.146 1.756
A2 cycles / f_cir, RF eta learned 1.556 3.362 1.653
'median' and 'p90' pool all 166 configurations, so hpcg (44 of them)
dominates. 'macro' is the mean of the six per-application medians.
--- paired Wilcoxon against the best reference ('t_cir = t_a2 * f_a2/f_cir (clock ratio)', n=166) ---
t_cir = t_a2 (no adjustment) wins 35/166 p= 5.24e-21 WORSE
median training Cirrus runtime wins 28/166 p= 7.92e-22 WORSE
t_a2 * constant speedup (fitted) wins 105/166 p= 0.957 WORSE
RF on A2 counters -> speedup wins 77/166 p= 0.442 WORSE
Ridge on A2 counters -> speedup wins 59/166 p= 1.31e-09 WORSE
RF on A2 counters -> log t_cir wins 75/166 p= 0.00587 WORSE
A2 cycles / f_cir, constant eta wins 100/166 p= 0.576 WORSE
A2 cycles / f_cir, RF eta wins 94/166 p= 0.993 better
--- the circular reference, on the same 166 configurations ---
CIRRUS cycles / f_cir, constant eta circular 1.129 3.601 1.475
This uses PAPI_TOT_CYC measured on Cirrus, i.e. it presupposes the
run it claims to predict. It is the ceiling the non-circular
predictors are being asked to approach, and the gap is the honest
cost of not having run on the target.
--- how far is the cycle count from being machine invariant? ---
cyc_cir / cyc_a2: median 0.651, IQR 0.451-1.195, sd(log10) 0.269
comd median 0.520 range 0.409- 1.148
hpcg median 1.666 range 0.914- 2.164
hpl median 0.418 range 0.192- 0.839
lulesh median 0.378 range 0.216- 0.693
minife median 0.474 range 0.149- 0.794
stream median 0.827 range 0.795- 1.508
--- per-application median error ---
t_cir = t_a2 (no adjustment) t_cir = t_a2 * f_a2/f_cir (clock ratio) median training Cirrus runtime t_a2 * constant speedup (fitted) RF on A2 counters -> speedup Ridge on A2 counters -> speedup RF on A2 counters -> log t_cir A2 cycles / f_cir, constant eta A2 cycles / f_cir, RF eta
app
comd 3.647 2.212 4.023 1.626 1.431 1.767 1.353 1.832 1.451
hpcg 1.182 1.395 69.327 2.803 2.786 3.918 6.128 2.403 2.810
hpl 4.131 2.505 4.592 1.842 1.297 1.609 1.499 1.830 1.256
lulesh 4.322 2.621 1.794 1.912 1.580 3.428 1.466 1.900 1.745
minife 2.736 1.846 5.329 1.582 1.302 1.511 1.358 1.480 1.338
stream 2.243 1.360 2.362 1.218 1.639 16.002 1.727 1.091 1.319
--- oracle diagnostics ---
oracle per-application constant speedup median 1.130 p90 1.669
(fitted on the test application itself, so not achievable; it
bounds what any method that only rescales t_a2 per application
could reach)
oracle single global constant speedup median 1.662 p90 2.259