🤖 AI Summary
This study addresses the persistent trade-off between behavioral interpretability and predictive performance in travel route and activity choice modeling by proposing a behaviorally disciplined hybrid modeling paradigm. We construct a unified framework integrating inverse reinforcement learning and imitation learning, revealing a shared soft Bellman structure across models while systematically incorporating constraint learning and large language models. This research establishes a theoretical linkage between route and activity choices, ensuring adherence to behavioral regularities and valid policy evaluation. Consequently, the proposed approach significantly enhances model scalability and prediction accuracy, achieving an organic unification of machine learning performance with robust behavioral identification capabilities.
📝 Abstract
Route and activity choice are connected levels of a common sequential mobility decision problem: activity choice determines what people do, where, and when, while route choice governs how they move between activities. This review develops a unified framework connecting transportation choice modeling with inverse reinforcement learning (IRL) and imitation learning (IL). Under explicit assumptions, recursive logit, logit dynamic discrete choice, and maximum-entropy IRL share a soft Bellman representation, while trajectory occupancies and network flows satisfy related conservation laws. However, utility, reward, policy, occupancy, constraints, and observation errors remain different estimands with different behavioral and counterfactual interpretations. We review constrained and inverse-constrained learning, occupancy-ratio and DICE methods, incomplete and mixed-quality demonstrations, graph and sequence learning, transfer, data fusion, multi-agent choice, and large language models. Our central message is that machine learning adds the greatest value when embedded within a behaviorally disciplined framework: exact transitions enforce feasibility, structured rewards preserve interpretable trade-offs, observation models address heterogeneous data sources, and network or equilibrium solvers produce coherent system outcomes. Such hybrid models can improve scalability and prediction without sacrificing behavioral identification or policy relevance.