An analysis of machine learning approaches for enhancing decision-making in complex discrete choice tasks

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Traditional parametric discrete choice models struggle to capture complex decision rules under individual heterogeneity, particularly in modeling policy preferences. This study systematically evaluates four machine learning approaches—multinomial logistic regression, generalized additive models, Siamese neural networks, and Gaussian processes—across five behavioral and social science discrete choice tasks. Using both Monte Carlo simulations and real-world energy policy preference data, models are compared via Bayesian Information Criterion (BIC) and predictive accuracy. Results show that semi-parametric and non-parametric models consistently outperform parametric ones; increasing training sample size and choice rule determinism improves performance by 6%–96% and 0%–55%, respectively. On empirical data, the Siamese neural network achieves the best fit (BIC = 13.351). These findings underscore that model selection should be guided by task characteristics and highlight both the promise and limitations of data-driven methods in policy-oriented discrete choice modeling.
📝 Abstract
Discrete choice modeling is a common tool used for preference elicitation during policy-making, but this is typically done through parametric models. Machine learning can push the boundaries of discrete choice modeling for policy-based preference elicitation by adopting a data-driven approach or learning individual preferences. However, there is limited knowledge of how well machine learning methods can estimate individual discrete choice rules under individual heterogeneity, especially in the context of challenges often experienced during preference elicitation. This study evaluates four machine learning models (multinomial logistic regression, generalized additive model, twinned neural network, and Gaussian process) with respect to their capacity to learn and predict five choice rules that are important in the behavioral and social sciences (linear strong utility, monotonic strong utility, ideal point, lexicographic semiorder, and multiattribute linear ballistic accumulator). Monte Carlo experiments were performed to assess model performance when increasing a) the number of attributes in the choice alternatives, b) the number of training choice sets, and c) the choice rule's determinism. The simulation results demonstrated that semi-parametric and non-parametric models generally outperform parametric models across all choice rules and experimental contexts. Model performance also generally improves by 6% to 96% and 0% to 55%, respectively, with an increase in training choice sets and choice rule determinism. A case study using real energy policy preference data was also conducted, where TNN performed best with a BIC of 13.351. This work demonstrated the viability and limitations of semi-parametric and non-parametric models in the context of policy-centric discrete choice modeling and showed how the choice task context should drive model selection.
Problem

Research questions and friction points this paper is trying to address.

discrete choice modeling
machine learning
individual heterogeneity
preference elicitation
choice rules
Innovation

Methods, ideas, or system contributions that make the work stand out.

machine learning
discrete choice modeling
individual heterogeneity
non-parametric models
preference elicitation
S
Sheng Lun Christine Cao
Department of Electrical and Software Engineering, University of Calgary, Calgary, AB, Canada
D
Destenie Nock
Department of Engineering and Public Policy, Carnegie Mellon University, Pittsburgh, PA, USA; Department of Civil Engineering, Carnegie Mellon University, Pittsburgh, PA, USA
A
Alex Davis
Independent Decision Science Consultant, Pittsburgh, PA, USA