On the Limits of Machine-Learned Ranking for Modern Microarchitectural Policies

šŸ“… 2026-08-02
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
This study systematically evaluates the reliability of four machine learning–based ranking models—NeuroScalar, SimNet, Concorde, and OneDSE—in ordering hardware configurations at the program phase level for microarchitectural design space exploration. Across structural parameter and behavioral policy scenarios, the analysis—integrating cycle-accurate simulation, Bayesian accuracy assessment, and information-theoretic methods—reveals, for the first time, that a substantial fraction (22.4%) of program windows exhibit counterintuitive rankings in structural settings, and inter-model consistency remains low (23.3%–39.9%). In behavioral policy scenarios, most models fail to surpass a featureless baseline, with the best achieving only a 2.1-percentage-point improvement. The work further establishes a theoretical upper bound on ranking accuracy when critical microarchitectural states are unobservable, demonstrating inherent limitations of instruction-stream–based approaches.
šŸ“ Abstract
Machine-learning predictors estimate processor performance far faster than cycle-level simulation. For design-space exploration, however, the valuable test is not merely reproducing the usual hardware ordering, but identifying how different hardware configurations rank on individual program phases. We evaluate four ML-predictors in two design regimes: \emph{Structural Parameters} (SP), varying hardware resources such as issue width, ROB size, and cache capacity; and \emph{Behavioral Policies} (BP), varying prefetching and replacement algorithms. In the SP regime, aggregate ranking is strong, yet counter-intuitive windows(CIW)---where the configuration expected to be slower is faster---constitute $22.4\%$ of non-tied windows across five pairs with a clear architectural prior. CIW match across these pairs is only $23.3$--$39.9\%$; every point estimate is below the $50\%$ random strict-ordering reference. The BP regime presents a different failure: ground-truth ties cover $37.8\%$ of pair-windows, most strict pairs have margins of only a few cycles, and no model family reliably beats a feature-free majority baseline. NeuroScalar and SimNet fall below that baseline, Concorde is statistically tied with it, and the best selected OneDSE head improves by only $2.1$ percentage points. Accuracy rises mainly at large margins. We further show that this failure is not a matter of model capacity: an information-theoretic analysis reveals that when ranking outcomes depend on hidden microarchitectural state absent from the instruction stream, no trace-based predictor can exceed the Bayes accuracy determined by observable inputs alone. Thus high cycle or aggregate ranking accuracy can reflect mastery of easy, high-margin cases while missing the local reversals that carry the most architectural insight and for which cycle-level simulation remains indispensable.
Problem

Research questions and friction points this paper is trying to address.

machine-learned ranking
microarchitectural policies
design-space exploration
performance prediction
cycle-level simulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

machine-learned ranking
microarchitectural policies
design-space exploration
information-theoretic limits
cycle-level simulation
šŸ”Ž Similar Papers
No similar papers found.
Y
Yanxin Zhang
NVIDIA
S
Shayne Wadle
University of Wisconsin–Madison
Y
Yuxuan Xiong
NVIDIA
Z
Zheyu Fu
NVIDIA
T
Trivikram Krishnamurthy
NVIDIA
K
Karu Sankaralingam
University of Wisconsin–Madison