Score
Using model selection and information-criterion methods to choose latent dimensionality or specification (and to compare models), and defining quantitative notions of accumulated evidence or 'specification debt' against deployed models.
This study addresses the challenge of reliably identifying key predictive variables in regression analysis to uncover underlying scientific mechanisms. We systematically evaluate variable selection strategies—including exhaustive search, greedy search, LASSO path, LASSO cross-validation, and random search—combined with AIC and BIC criteria, across linear and generalized linear models under varying model space sizes (small to high-dimensional). Our empirical comparison on simulated data reveals that BIC paired with exhaustive search (for small spaces) or random search (for large spaces) consistently achieves the highest true positive rate, lowest false discovery rate, and superior stability and reproducibility compared to all other combinations. These findings establish a theoretically grounded yet practically feasible paradigm for high-dimensional variable selection, offering robust, reproducible modeling guidance for scientific inference.
This study addresses the limitation of existing information criteria in structural equation modeling, which fail to explicitly leverage latent variable structures, thereby constraining model selection performance. To overcome this, the authors propose two novel information criteria based on the complete-data likelihood, uniquely integrating complete-data likelihood with importance sampling to explicitly incorporate latent variable structure into criterion design. Within the Gaussian structural equation modeling framework, the approach accurately estimates latent variables, thereby effectively recovering dependencies among observed variables. Empirical evaluations demonstrate that the proposed criteria exhibit robust performance across diverse latent structures and sample conditions, significantly outperforming existing methods—particularly when latent variable estimation is accurate.
Existing asymptotic optimality proofs for model selection criteria rely heavily on restrictive assumptions—specific model classes, estimation methods, and data structures—limiting their theoretical validity and practical applicability. Method: We develop a unified asymptotic theory framework that relaxes these constraints, accommodating diverse models (e.g., linear regression, quantile regression, penalized regression), heterogeneous data structures (independent, dependent, and high-dimensional), and broad estimation paradigms (maximum likelihood, generalized method of moments, linear smoothers). Contribution/Results: Within this general setting, we establish, for the first time, rigorous asymptotic optimality of canonical criteria—including AIC, BIC, and cross-validation—under minimal regularity conditions. Our results substantially broaden the theoretical scope of these criteria, enabling principled model selection in complex, real-world scenarios with dependent or high-dimensional data. The framework provides a robust, general foundation for statistical inference and model choice beyond conventional parametric and i.i.d. settings.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
This study addresses the challenge of adaptively determining when to reset a model’s structure under lightweight parameter update strategies to balance predictive accuracy, computational cost, and stability. The authors propose a “model specification debt” mechanism that accumulates evidence—such as prediction score discrepancies, stacked weights, or calibration diagnostics—to formulate a cost-sensitive trigger rule for model resetting. This framework generalizes fixed-interval updating as a special case and enables flexible deployment in open environments. Evaluated on the M4 dataset, the approach achieves predictive accuracy comparable to full retraining while consuming only 28% of the computation time, significantly reducing instability. It consistently matches or outperforms fixed-update strategies across diverse scenarios and offers dynamic, evidence-driven adaptation capabilities.
This work addresses the failure of traditional penalized methods in non-nested model selection by proposing a maximum likelihood-based criterion that selects the model with the highest likelihood without favoring simpler models or relying on explicit penalties for model complexity. Under the assumption that all candidate models are equally important a priori, the method directly chooses the model maximizing the likelihood and is theoretically shown to be statistically consistent. Both theoretical analysis and empirical experiments demonstrate that the proposed criterion not only achieves consistency in non-nested settings but also outperforms existing penalized model selection approaches in terms of selection accuracy and overall performance.
Traditional Bayesian modeling relies on model selection to balance complexity and generalization, yet this approach often compromises predictive performance in small-sample settings. This work proposes a “predictive consistency prior” that maintains stability in the prior predictive distribution as model complexity increases, thereby circumventing explicit model selection. By shifting the modeling focus from parameter sparsity to constructing reasonable and stable priors in predictive space, the method reveals that the perceived necessity of model selection fundamentally stems from inadequate prior specification. The authors implement this prior in Bayesian linear and logistic regression, forward variable selection, and nonlinear models, demonstrating through numerical experiments that flexible models equipped with the predictive consistency prior match or even outperform carefully selected simpler models in out-of-sample prediction across a range of tasks.
This study addresses the persistent challenge in econometrics of underfitting or overfitting in regression and vector autoregressive (VAR) models arising from ambiguous determination of the number of underlying states. To resolve this, the paper proposes a four-stage model specification framework grounded in combinatorial mathematics. The framework systematically enumerates feasible model configurations at each stage, quantifies—for the first time—the size of the latent state space, and integrates information criteria such as AIC and BIC to construct a comprehensive search mechanism across the entire specification space. By doing so, it fills a critical theoretical gap concerning the assessment of state dimensionality in correctly specified models, provides a rigorous foundation for bounding model complexity, exposes limitations of prevailing selection approaches, and ultimately advances more accurate and robust econometric modeling practices.
This work proposes the Focused Information Criterion (FIC), a model selection framework that departs from the conventional pursuit of global optimality by tailoring model choice to a user-specified target parameter or risk function—such as mean squared error—thereby prioritizing relevance to the inferential goal at hand. Unlike traditional criteria, FIC dynamically selects the model that minimizes the estimated risk for the quantity of interest, accommodating diverse modeling paradigms including linear, nonparametric, quantile regression, and graphical models. The approach further extends to high-dimensional, longitudinal, survival, and time series data through integration with regularization techniques, Bayesian estimation, and targeted risk optimization. This yields a unified and flexible framework for “goal-oriented” optimal modeling, where the best model is defined not universally but relative to the specific objective of the analysis.
This study addresses the inconsistency in model selection for multivariate probabilistic time series forecasting, where different summary statistics—such as mean, median, or average rank—often yield conflicting decisions. By leveraging proper scoring rules, the authors systematically analyze the distributional properties of model scores on test sets and demonstrate that skewness in these distributions is the primary cause of such discrepancies. Theoretical analysis and empirical evidence show that only the mean score consistently identifies the optimal model under short test sets, while average rank, though invariant to prediction scaling, is highly sensitive to distributional skew. Large-scale experiments on intermittent demand datasets like M5 confirm the critical influence of test set length on selection consistency, offering both theoretical grounding and practical guidance for robust model evaluation.