Opens a larger view. Escape closes it.

hardware-counters

literature_corpus.md

Citation corpus for the literature survey

Bibliography for the literature survey. Bibliographic details (authors, titles, venues, years, DOIs and arXiv identifiers) were taken from Crossref, OpenAlex, arXiv and Semantic Scholar rather than transcribed by hand, so they should be accurate as given.

The short note after each entry records why the paper is relevant to this project. Those notes are based on abstracts and, where read, the papers themselves; they are not a substitute for reading the work.

Around forty papers were retained after screening several hundred search results across the four sources. Entries are grouped by topic below.


1. Analytical / mechanistic models

  • Williams, S.; Waterman, A.; Patterson, D. Roofline: an insightful visual performance model for multicore architectures. Communications of the ACM 52(4):65–76, 2009. doi:10.1145/1498765.1498785 — bound-based model: attainable FLOP/s = min(peak, bandwidth x arithmetic intensity).
  • Williams, S.; Patterson, D.; Oliker, L.; Shalf, J.; Yelick, K. The Roofline model: a pedagogical tool for program analysis and optimization. IEEE Hot Chips 20, 2008. doi:10.1109/hotchips.2008.7476531 — earlier Roofline venue.
  • Ilić, A.; Pratas, F.; Sousa, L. Cache-aware Roofline model: upgrading the loft. IEEE Computer Architecture Letters 13(1):21–24, 2014. doi:10.1109/l-ca.2013.6 — adds cache levels; validated using hardware counters, curve fitness > 90%.
  • Ilić, A.; Pratas, F.; Sousa, L. Beyond the Roofline: cache-aware power and energy-efficiency modeling for multi-cores. IEEE Trans. Computers 66(1):52–58, 2017. doi:10.1109/tc.2016.2582151 — power/energy bounds split into cores, uncore and package domains; validated with hardware counters and on-chip power monitoring on Intel 3770K.
  • Lo, Y.J.; Williams, S.; Van Straalen, B.; Ligocki, T.J.; Cordery, M.J.; Wright, N.J. et al. Roofline Model Toolkit: a practical tool for architectural and program analysis. LNCS (PMBS), 2015. doi:10.1007/978-3-319-17248-4_7
  • Hager, G.; Treibig, J.; Habich, J.; Wellein, G. Exploring performance and power properties of modern multicore chips via simple machine models. Concurrency and Computation: Practice and Experience 28(2):189–210, 2013. doi:10.1002/cpe.3180 ; arXiv:1208.2908 — ECM model + phenomenological power model; predicts saturation behaviour Roofline cannot.
  • Stengel, H.; Treibig, J.; Hager, G.; Wellein, G. Quantifying performance bottlenecks of stencil computations using the Execution-Cache-Memory model. ICS ‘15, pp. 207–216. doi:10.1145/2751205.2751240 ; arXiv:1410.5010 — refined ECM; single-core and scaling predictions; explicit comparison to Roofline.
  • Hofmann, J.; Eitzinger, J.; Fey, D. Execution-Cache-Memory performance model: introduction and validation. arXiv:1509.03118, 2015.
  • Hofmann, J.; Alappat, C.L.; Hager, G.; Fey, D.; Wellein, G. Bridging the architecture gap: abstracting performance-relevant properties of modern server processors. Supercomputing Frontiers and Innovations 7(2), 2020. doi:10.14529/jsfi200204
  • Hammer, J.; Eitzinger, J.; Hager, G.; Wellein, G. Kerncraft: a tool for analytic performance modeling of loop kernels. Tools for HPC 2016, LNCS, 2017. doi:10.1007/978-3-319-56702-0_1 ; arXiv:1702.04653 — automates ECM/ Roofline from source.
  • Eyerman, S.; Eeckhout, L.; Karkhanis, T.; Smith, J.E. A performance counter architecture for computing accurate CPI components. ASPLOS ‘06 / ACM SIGPLAN Notices 41(11):175–184, 2006. doi:10.1145/1168918.1168880 — CPI stacks via interval analysis; the canonical decomposition of cycles into baseline + miss-event components.
  • Eyerman, S.; Eeckhout, L.; Karkhanis, T.; Smith, J.E. A top-down approach to architecting CPI component performance counters. IEEE Micro 27(1), 2007. doi:10.1109/mm.2007.3
  • Eyerman, S.; Hoste, K.; Eeckhout, L. Mechanistic-empirical processor performance modeling for constructing CPI stacks on real hardware. ISPASS 2011, pp. 216–226. doi:10.1109/ispass.2011.5762738 — grey-box: mechanistic functional form, regression-fitted coefficients; Pentium 4 / Core 2 / Core i7 with SPEC CPU2000+2006; average prediction error 9–13%. Closest published analogue to our eta framing.
  • Yasin, A. A top-down method for performance analysis and counters architecture. ISPASS 2014, pp. 35–44. doi:10.1109/ispass.2014.6844459 — Top-Down Microarchitecture Analysis (TMA); hierarchical slot-based bottleneck attribution using ~8 counters. 322 citations.
  • Culler, D.; Karp, R.; Patterson, D.; Sahay, A.; Schauser, K.E.; Santos, E.; Subramonian, R.; von Eicken, T. LogP: towards a realistic model of parallel computation. PPoPP ‘93. doi:10.1145/155332.155333
  • Alexandrov, A.; Ionescu, M.F.; Schauser, K.E.; Scheiman, C. LogGP: incorporating long messages into the LogP model for parallel computation. J. Parallel and Distributed Computing 44(1):71–79, 1997. doi:10.1006/jpdc.1997.1346
  • Amdahl, G.M. Validity of the single processor approach to achieving large scale computing capabilities. AFIPS Spring Joint Computer Conf. 1967, p. 483. doi:10.1145/1465482.1465560
  • Gustafson, J.L. Reevaluating Amdahl’s law. Communications of the ACM 31(5):532–533, 1988. doi:10.1145/42411.42415
  • Afzal, A.; Hager, G.; Wellein, G. Analytic roofline modeling and energy analysis of the LULESH proxy application on multi-core clusters. Int. J. High Performance Computing Applications 40(1):123–141, 2025. doi:10.1177/10943420251363711 ; arXiv:2412.08792 — Roofline per hot spot on Ice Lake / Sapphire Rapids, validated against hardware counters; also power, energy-to-solution and EDP. Directly overlaps our LULESH + RAPL setup.

2. Hardware-counter-based prediction and ML

  • Browne, S.; Dongarra, J.; Garner, N.; Ho, G.; Mucci, P. A portable programming interface for performance evaluation on modern processors. Int. J. High Performance Computing Applications 14(3):189–204, 2000. doi:10.1177/109434200001400303 — PAPI.
  • Ding, N.; Lee, V.W.; Xue, W.; Zheng, W. APMT: an automatic hardware counter-based performance modeling tool for HPC applications. CCF Trans. High Performance Computing 2:135–148, 2020. doi:10.1007/s42514-020-00035-8 — counter-assisted kernel profiling with an optional CPI refinement framework; ~15% average error, 3% overhead; 25–52% better than analytic and empirical baselines in strong scaling. The single closest prior work to this project.
  • Malakar, P.; Balaprakash, P.; Vishwanath, V.; Morozov, V.; Kumaran, K. Benchmarking machine learning methods for performance modeling of scientific applications. PMBS @ SC 2018, pp. 33–44. doi:10.1109/pmbs.2018.8641686 — eleven ML methods, four applications, four leadership platforms; bagging/ boosting/DNN reach median R² > 0.95; studies training-set size, transfer learning and extrapolation explicitly.
  • Sun, J.; Sun, G.; Zhan, S.; Zhang, J.; Chen, Y. Automated performance modeling of HPC applications using machine learning. IEEE Trans. Computers 69(5):749–763, 2020. doi:10.1109/tc.2020.2964767 — random forest on domain-independent runtime features (variable values, branch/loop/MPI counters); Graph500, GalaxSee, SMG2000 on three systems; <20% mean error; transfer learning to a new platform.
  • Sun, J.; Zhan, S.; Sun, G.; Chen, Y. Automated performance modeling based on runtime feature detection and machine learning. ISPA/IUCC 2017. doi:10.1109/ispa/iucc.2017.00115
  • Orteu Aubach, J.; Banchelli, F.; Clascà Ramírez, M.; Garcia-Gasulla, M. Heuristic-based merging of HPC traces to extend hardware counter coverage. arXiv:2605.15832, 2026 — directly addresses our 5-counter limit: merges traces from multiple runs each with a different counter set, matching computation bursts via MPI structure/timing, to build a unified wide-feature dataset for ML performance prediction without multiplexing. Validated on MareNostrum5.
  • Gritz, M.; Silva, G.; Klôh, V.; Schulze, B.; Ferro, M. Towards an autonomous framework for HPC optimization: a study of performance prediction using hardware counters and machine learning. SPOLM 2019/2020. doi:10.5151/spolm2019-196
  • Das, S.; Werner, J.; Antonakakis, M.; Polychronakis, M.; Monrose, F. SoK: the challenges, pitfalls, and perils of using hardware performance counters for security. IEEE S&P 2019. doi:10.1109/sp.2019.00021 — counter non-determinism, overcounting and portability caveats.
  • Wang, Y.; Lee, V.; Wei, G.-Y.; Brooks, D. Predicting new workload or CPU performance by analyzing public datasets. ACM TACO 15(4):1–21, 2018. doi:10.1145/3284127 — DNN over SPEC CPU2006 / Geekbench 3 public results; predicts new processors and new workloads; quantifies benchmark self-similarity.

3. Cross-application / cross-platform generalisation

  • Ardalani, N.; Lestourgeon, C.; Sankaralingam, K.; Zhu, X. Cross-architecture performance prediction (XAPP) using CPU code to predict GPU performance. MICRO-48, 2015, pp. 725–737. doi:10.1145/2830772.2830780 — ML from single-threaded CPU features to GPU speedup; explicitly a cross-target generalisation problem.
  • Yang, L.T.; Ma, X.; Mueller, F. Cross-platform performance prediction of parallel applications using partial execution. SC ‘05, 2005. doi:10.1109/sc.2005.20 — relative performance between two platforms observed from short partial executions; no modelling or simulation.
  • Mahdavi, K. A hybrid machine learning method for cross-platform performance prediction of parallel applications. ICPP 2024, pp. 669–678. doi:10.1145/3673038.3673059 — predicts ratios between platforms from brief partial executions on a reference platform; “Ensemble Cluster Classify Regress”.
  • Marathe, A.; Anirudh, R.; Jain, N.; Bhatele, A.; Thiagarajan, J.; Kailkhura, B.; Yeom, J.-S.; Rountree, B.; Gamblin, T. Performance modeling under resource constraints using deep transfer learning. SC ‘17. doi:10.1145/3126908.3126969 — transfers from cheap small-scale observations to a large target scale; identifies best configurations with as few as 1% of target-scale observations.
  • Jamshidi, P.; Siegmund, N.; Velez, M.; Kästner, C.; Patel, A.; Agarwal, Y. Transfer learning for performance modeling of configurable systems: an exploratory analysis. ASE 2017, pp. 497–508. doi:10.1109/ase.2017.8115661 — when transfer works: small environment changes admit a linear correction, severe changes transfer only sampling structure. Directly relevant to our “unseen application = severe distribution shift” claim.
  • Ferrerón, A.; Jagtap, R.; Bischoff, S.; Rusitoru, R. Crossing the architectural barrier: evaluating representative regions of parallel HPC applications. ISPASS 2017, pp. 109–120. doi:10.1109/ispass.2017.7975275 ; arXiv:1803.09584 — BarrierPoint across Intel and ARM; error below 2.3% for cycles and instructions, 178x simulation-time reduction. Also IISWC 2016 version, doi:10.1109/iiswc.2016.7581284, error < 3.3%.
  • Barnes, B.J.; Rountree, B.; Lowenthal, D.K.; Reeves, J.; de Supinski, B.; Schulz, M. A regression-based approach to scalability prediction. ICS ‘08, pp. 368–377. doi:10.1145/1375527.1375580
  • Barnes, B.; Garren, J.; Lowenthal, D.K.; Reeves, J.; de Supinski, B.R.; Schulz, M.; Rountree, B. Using focused regression for accurate time-constrained scaling of scientific applications. IPDPS 2010. doi:10.1109/ipdps.2010.5470431 — grey-box regression from a small set of training runs; median prediction error < 13% across seven applications.
  • Hoefler, T.; Belli, R. Scientific benchmarking of parallel computing systems. SC ‘15. doi:10.1145/2807591.2807644 — stratified sample of 120 papers; the standard reference for statistically sound HPC reporting (non-parametric statistics, reporting variability). Justifies our Wilcoxon protocol.

4. Proxy / mini-apps as predictors of full applications

  • Heroux, M.A.; Crozier, P.; Thornquist, H.; Numrich, R.; Williams, A.; Edwards, H.C.; Keiter, E. et al. Improving performance via mini-applications. Sandia National Laboratories Technical Report SAND2009-5574, 2009. doi:10.2172/993908 — the Mantevo report; origin of miniFE et al.
  • Barrett, R.F.; Crozier, P.S.; Doerfler, D.W.; Heroux, M.A.; Lin, P.T.; Thornquist, H.K.; Trucano, T.G.; Vaughan, C.T. Assessing the role of mini-applications in predicting key performance characteristics of scientific and engineering applications. J. Parallel and Distributed Computing 75:107–122, 2015. doi:10.1016/j.jpdc.2014.09.006 — the canonical fidelity study.
  • Barrett, R.; Crozier, P.; Doerfler, D.; Hammond, S.; Heroux, M.; Lin, P. et al. Poster: assessing the predictive capabilities of mini-applications. SC Companion 2012. doi:10.1109/sc.companion.2012.169
  • Aaziz, O.; Cook, J. (Jeanine); Cook, J. (Jonathan); Juedeman, T.; Richards, D.F.; Vaughan, C. A methodology for characterizing the correspondence between real and proxy applications. IEEE Cluster 2018, pp. 190–200. doi:10.1109/cluster.2018.00037 — LDMS hardware-counter + mpiP data on two platforms, four ECP proxy/parent pairs; finds proxies representative for compute and memory behaviour.
  • Aaziz, O.; Vaughan, C.; Cook, J.; Cook, J.; Kuehn, J.A.; Richards, D.F. Fine-grained analysis of communication similarity between real and proxy applications. PMBS @ SC 2019, pp. 93–102. doi:10.1109/pmbs49563.2019.00016 — identifies at least one parent/proxy mismatch; the negative-result counterweight to Aaziz 2018.
  • Richards, D.; Aaziz, O.; Cook, J.; Kuehn, J.; Moore, S.; Pruitt, D. et al. Quantitative performance assessment of proxy apps and parents. ECP Proxy App Project Milestone reports: ADCD-504-9, 2020, doi:10.2172/1617284; ADCD-504-11, 2021, doi:10.2172/1860797; 2022, doi:10.2172/2432204.
  • Richards, D.F.; Aaziz, O.; Alexeev, Y.; Balakrishnan, R.; Cook, J.; Finkel, H. et al. FY20 / FY21 proxy app suite release. ECP milestone reports, doi:10.2172/1860724 (FY20), doi:10.2172/1860683 (FY21) — source of the ECP suite containing miniFE, LULESH-family and CoMD-family codes.
  • Karlin, I.; Bhatele, A.; Keasler, J.; Chamberlain, B.L.; Cohen, J.; DeVito, Z.; Haque, R.; Laney, D.; Luke, E.; Wang, F.; Richards, D.; Schulz, M.; Still, C. Exploring traditional and emerging parallel programming models using a proxy application. IPDPS 2013, pp. 919–932. doi:10.1109/ipdps.2013.115 — LULESH.
  • Pearce, O.; Ahmed, H.; Larsen, R.W.; Pirkelbauer, P.; Richards, D.F. Exploring dynamic load imbalance solutions with the CoMD proxy application. Future Generation Computer Systems 92:920–932, 2019. doi:10.1016/j.future.2017.12.010
  • Owenson, A.M.B.; Wright, S.A.; Bunt, R.; Ho, Y.K.; Street, M.J.; Jarvis, S.A. An unstructured CFD mini-application for the performance prediction of a production CFD code. Concurrency and Computation: Practice and Experience 32(10), 2019. doi:10.1002/cpe.5443 — MG-CFD proxy for Rolls-Royce HYDRA; predicts the parent code’s strong scaling with mean error 9.2%. The cleanest quantitative demonstration that a proxy can stand in for a real code.
  • Matsuoka, S.; Domke, J.; Wahib, M.; Drozd, A.; Chien, A.A.; Bair, R.A.; Vetter, J.S.; Shalf, J. Preparing for the future — rethinking proxy applications. Computing in Science & Engineering 24(2):85–90, 2022. doi:10.1109/mcse.2022.3153105 ; arXiv:2204.07336 — argues proxy suites induce rigidity and are a poor fit for a heterogeneous future.
  • McKinsey, M.; Brink, S.; Pearce, O. On similarity of computational kernels in our codes and proxies. arXiv:2605.06968, 2026 — automated hardware-usage-based similarity metrics between proxy and parent kernels (Kripke vs RAJA Performance Suite), CPU and GPU.
  • Dongarra, J.; Heroux, M.A.; Łuszczek, P. High-performance conjugate-gradient benchmark: a new metric for ranking high-performance computing systems. Int. J. High Performance Computing Applications 30(1):3–10, 2015. doi:10.1177/1094342015593158 — HPCG.
  • Dongarra, J.J.; Łuszczek, P.; Petitet, A. The LINPACK benchmark: past, present and future. Concurrency and Computation: Practice and Experience 15(9):803–820, 2003. doi:10.1002/cpe.728 — HPL.

5. Small-sample ML in systems research

  • Ritter, M.; Calotoiu, A.; Rinke, S.; Reimann, T.; Hoefler, T.; Wolf, F. Learning cost-effective sampling strategies for empirical performance modeling. IPDPS 2020, pp. 884–895. doi:10.1109/ipdps47924.2020.00095 — polynomial rather than exponential number of experiments per parameter; 85% cost reduction retaining 92% of model accuracy.
  • Ritter, M.; Naumann, B.; Calotoiu, A.; Rinke, S.; Reimann, T.; Hoefler, T.; Wolf, F. Cost-effective empirical performance modeling. IEEE Trans. Parallel and Distributed Systems 37(3):575–592, 2026. doi:10.1109/tpds.2025.3646119 — Gaussian-process-regression-guided point selection, per modelling task.
  • Ritter, M.; Geis, A.B.U.; Wehrstein, J.; Calotoiu, A.; Reimann, T.; Hoefler, T.; Wolf, F. Noise-resilient empirical performance modeling with deep neural networks. IPDPS 2021, pp. 23–34. doi:10.1109/ipdps49936.2021.00012 — +25% model accuracy at high noise, +15% predictive power.
  • Ha, H.; Zhang, H. DeepPerf: performance prediction for configurable software with deep sparse neural network. ICSE 2019, pp. 1095–1106. doi:10.1109/icse.2019.00113 — L1-sparse FNN with automated hyperparameter search; higher accuracy with less training data than prior work on eleven real-world datasets. The canonical “small-n deep model” result.
  • Kaltenecker, C.; Grebhahn, A.; Siegmund, N.; Guo, J.; Apel, S. Distance-based sampling of software configuration spaces. ICSE 2019, pp. 1084–1094. doi:10.1109/icse.2019.00112
  • Kaltenecker, C.; Grebhahn, A.; Siegmund, N.; Apel, S. The interplay of sampling and machine learning for software performance prediction. IEEE Software 37(4):58–66, 2020. doi:10.1109/ms.2020.2987024
  • De Sensi, D. Predicting performance and power consumption of parallel applications. PDP 2016, pp. 200–207. doi:10.1109/pdp.2016.41 — multiple linear regression over cores x frequency; 96% average accuracy from 1% of the configuration space on PARSEC.

6. Extrapolation vs interpolation

  • Calotoiu, A.; Hoefler, T.; Poke, M.; Wolf, F. Using automated performance modeling to find scalability bugs in complex codes. SC ‘13. doi:10.1145/2503210.2503277 — Extra-P: per-kernel empirical scaling models from small-scale runs, extrapolated to large core counts.
  • Calotoiu, A.; Beckinsale, D.; Earl, C.W.; Hoefler, T.; Karlin, I.; Schulz, M.; Wolf, F. Fast multi-parameter performance modeling. IEEE Cluster 2016, pp. 172–181. doi:10.1109/cluster.2016.57
  • Calotoiu, A.; Copik, M.; Hoefler, T.; Ritter, M.; Shudler, S.; Wolf, F. ExtraPeak: advanced automatic performance modeling for HPC applications. LNCSE, 2020, pp. 453–482. doi:10.1007/978-3-030-47956-5_15
  • Calotoiu, A.; Copik, M.; Czappa, F.; Geiss, A.; de Morais, G.; Ritter, M.; Shudler, S.; Hoefler, T.; Wolf, F. Extra-P — empirical performance modeling made easy. Frontiers in High Performance Computing 3, 2026. doi:10.3389/fhpcp.2025.1714042 — current overview: Performance Model Normal Form, sparse modelling, GPR, deep learning for noise, segmented modelling.
  • Shudler, S.; Calotoiu, A.; Hoefler, T.; Strube, A.; Wolf, F. Exascaling your library: will your implementation meet your expectations? ICS ‘15, pp. 165–175. doi:10.1145/2751205.2751216 — validates asymptotic scaling trends rather than point predictions.
  • Shudler, S.; Calotoiu, A.; Hoefler, T.; Wolf, F. Isoefficiency in practice: configuring and understanding the performance of task-based applications. PPoPP ‘17, pp. 131–143. doi:10.1145/3018743.3018770 — empirical isoefficiency function binding efficiency, core count and input size. The closest published relative of our dimensionless-efficiency framing.
  • Ritter, M.; Wolf, F. Extra-Deep: automated empirical performance modeling for distributed deep learning. SC ‘23 Workshops, pp. 1345–1356. doi:10.1145/3624062.3624204 — 93.6% average prediction accuracy, 94.9% profiling-time reduction.
  • Vetter, J.S.; McCracken, M.O. Statistical scalability analysis of communication operations in distributed applications. PPoPP ‘01. doi:10.1145/379539.379590

7. Energy / power modelling

  • Contreras, G.; Martonosi, M. Power prediction for Intel XScale processors using performance monitoring unit events. ISLPED ‘05. doi:10.1145/1077603.1077657 — linear counter-weighted power model; within 4% of measured average CPU power. The seminal counter-based power model.
  • Nagasaka, H.; Maruyama, N.; Nukada, A.; Endo, T.; Matsuoka, S. Statistical power modeling of GPU kernels using performance counters. Int. Green Computing Conf. 2010, pp. 115–122. doi:10.1109/greencomp.2010.5598315 — linear regression over GPU counters, 49 kernels, average error 4.7%; fails on kernels whose activity is uncounted (texture reads).
  • Huang, W.; Lefurgy, C.; Kuk, W.; Buyuktosunoglu, A.; Floyd, M.; Rajamani, K.; Allen-Ware, M.; Brock, B. Accurate fine-grained processor power proxies. MICRO-45, 2012, pp. 224–234. doi:10.1109/micro.2012.29 — per-core power proxy in production firmware; 1.8% mean unsigned error at 32 ms granularity.
  • Tiwari, A.; Laurenzano, M.A.; Carrington, L.; Snavely, A. Modeling power and energy usage of HPC kernels. IPDPSW 2012, pp. 990–998. doi:10.1109/ipdpsw.2012.121 — ANN CPU and DIMM power/energy models; average absolute error < 5.5% on MM, stencil, LU.
  • Hackenberg, D.; Schöne, R.; Ilsche, T.; Molka, D.; Schuchart, J.; Geyer, R. An energy efficiency feature survey of the Intel Haswell processor. IPDPSW 2015, pp. 896–904. doi:10.1109/ipdpsw.2015.70 — RAPL moves from modelling to actual measurement; also documents that opportunistic clocks vastly decrease performance predictability. Directly relevant to our observed 2.25 GHz-vs-effective-clock gap.
  • Khan, K.N.; Hirki, M.; Niemi, T.; Nurminen, J.K.; Ou, Z. RAPL in action: experiences in using RAPL for power measurements. ACM Trans. Modeling and Performance Evaluation of Computing Systems 3(2):1–26, 2018. doi:10.1145/3177754 — RAPL is highly correlated with wall power, accurate enough, negligible overhead. 269 citations. The standard justification for trusting our PACKAGE/PP0 readings.
  • Desrochers, S.; Paradis, C.; Weaver, V.M. A validation of DRAM RAPL power measurements. MEMSYS ‘16. doi:10.1145/2989081.2989088
  • Shahid, A.; Fahad, M.; Manumachu, R.R.; Lastovetsky, A. Energy predictive models of computing: theory, practical implications and experimental analysis on multicore processors. IEEE Access 9:63149–63172, 2021. doi:10.1109/access.2021.3075139 — a theory of PMC-based energy models; the additivity selection criterion improves linear model error from 31.2% to 18%.
  • Shahid, A.; Fahad, M.; Manumachu, R.R.; Lastovetsky, A. A comparative study of techniques for energy predictive modeling using performance monitoring counters on modern multicore CPUs. IEEE Access 8:143306–143332, 2020. doi:10.1109/access.2020.3013812 — theory-selected linear regression beats random forest by 5.09x and neural networks by 4.37x at platform level. Strong counterweight to “just use a tree ensemble”.
  • Shahid, A.; Fahad, M.; Manumachu, R.R.; Lastovetsky, A. Improving the accuracy of energy predictive models for multicore CPUs by combining utilization and performance events model variables. J. Parallel and Distributed Computing 151:38–51, 2021. doi:10.1016/j.jpdc.2021.01.007
  • Fahad, M.; Shahid, A.; Manumachu, R.R.; Lastovetsky, A. A comparative study of methods for measurement of energy of computing. Energies 12(11):2204, 2019. doi:10.3390/en12112204
  • Wu, X.; Taylor, V.; Cook, J.; Mucci, P.J. Using performance-power modeling to improve energy efficiency of HPC applications. Computer 49(10):20–29, 2016. doi:10.1109/mc.2016.311
  • Hofmann, J.; Hager, G.; Fey, D. On the accuracy and usefulness of analytic energy models for contemporary multicore processors. ISC 2018, LNCS. doi:10.1007/978-3-319-92040-5_2 ; arXiv:1803.01618
  • Fan, K.; Cosenza, B.; Juurlink, B. Accurate energy and performance prediction for frequency-scaled GPU kernels. Computation 8(2):37, 2020. doi:10.3390/computation8020037
  • Patki, T.; Lowenthal, D.K.; Sasidharan, A.; Maiterth, M.; Rountree, B.L.; Schulz, M.; de Supinski, B.R. Practical resource management in power-constrained, high performance computing. HPDC ‘15. doi:10.1145/2749246.2749262

Ancillary / tooling

  • Perelman, E.; Hamerly, G.; Van Biesbrouck, M.; Sherwood, T.; Calder, B. Using SimPoint for accurate and efficient simulation. SIGMETRICS ‘03. doi:10.1145/781027.781076
  • Arafa, Y.; Badawy, A.-H.A.; Chennupati, G.; Santhi, N.; Eidenbenz, S. PPT-GPU: scalable GPU performance modeling. IEEE Computer Architecture Letters 18(1):55–58, 2019. doi:10.1109/lca.2019.2904497 — within 10% of real devices, up to 450x faster than GPGPU-Sim.
  • Barai, A.; Arafa, Y.; Badawy, A.-H.; Chennupati, G.; Santhi, N.; Eidenbenz, S. PPT-Multicore: performance prediction of OpenMP applications using reuse profiles and analytical modeling. J. Supercomputing 78:2354–2385, 2021. doi:10.1007/s11227-021-03949-4

Searched for and not found

  • No published work was found that holds out an entire application as the test fold for a hardware-counter-based runtime model on a single node. Queries via Crossref and OpenAlex for “leave-one-out application holdout performance prediction” returned only unrelated (biomedical) leave-one-out literature. This appears to be a genuine gap.
  • No published work was found reporting PP0/PACKAGE energy ratio as a dominant predictive feature for runtime. The nearest results are Ilić et al. 2017 (core/uncore/package power decomposition, but as an analytic bound, not an ML feature) and the Shahid/Lastovetsky series (counters -> energy, i.e. the opposite direction).
  • McCalpin’s STREAM technical report could not be resolved to a DOI via Crossref or OpenAlex and is cited in the survey as a technical report, by URL.