hardware-counters
retrain2_summary.txt
Corrected evaluation
==============================================================================
F2: merged 191 configurations (was 815 pseudo-rows)
feature matrix now 97.4% populated (was ~42% on per-set rows)
191 configs, 21 features, 8 applications
--- LOAO (train on 7 apps, predict the 8th) ---
Constant (median log-eta) [F1] median 1.065 p90 1.532
Random Forest median 1.054 p90 1.390
GBM quantile(0.5) [F5] median 1.062 p90 1.367
MLP (32,16) a=1 [tuned on test] median 1.076 p90 1.319
MLP nested selection [F3] median 1.063 p90 1.311
inner loop picked: comd:(32, 16)a1.0, gromacs:(32, 16)a1.0, hpcg:(16,)a3.0, hpl:(16,)a3.0, lulesh:(64, 32)a1.0, minife:(16,)a3.0, openfoam:(16,)a3.0, stream:(16,)a3.0
MLP shrunk 0.5 to median [F5] median 1.068 p90 1.339
MLP + RF ensemble median 1.060 p90 1.270
--- F6: paired Wilcoxon at CONFIG granularity (n=191) ---
Random Forest vs constant: wins 125/191 p=0.0002162 (better)
GBM quantile(0.5) [F5] vs constant: wins 117/191 p=7.847e-05 (better)
MLP (32,16) a=1 [tuned on test] vs constant: wins 100/191 p=0.2897 (WORSE)
MLP nested selection [F3] vs constant: wins 94/191 p=0.3079 (better)
MLP shrunk 0.5 to median [F5] vs constant: wins 125/191 p=6.025e-06 (WORSE)
MLP + RF ensemble vs constant: wins 106/191 p=0.001438 (better)
--- per-application median (LOAO) ---
Constant (median log-eta) [F1] Random Forest GBM quantile(0.5) [F5] MLP (32,16) a=1 [tuned on test] MLP nested selection [F3] MLP shrunk 0.5 to median [F5] MLP + RF ensemble
app
comd 1.046 1.026 1.030 1.182 1.182 1.099 1.092
gromacs 1.051 1.007 1.018 1.017 1.017 1.033 1.009
hpcg 1.048 1.117 1.063 1.056 1.071 1.054 1.080
hpl 1.078 1.035 1.049 1.028 1.041 1.050 1.030
lulesh 1.175 1.048 1.102 1.026 1.071 1.094 1.036
minife 1.372 1.061 1.098 1.120 1.141 1.258 1.059
openfoam 1.068 1.075 1.097 1.054 1.053 1.059 1.068
stream 1.051 1.095 1.072 1.065 1.051 1.012 1.046