ALIGN-HOLD: Experience Alignment for Real-Time Hold Control in Large-Scale Ride-Hailing Matching at DiDi

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出ALIGN-HOLD框架,通过学习市场隐含偏好解决大规模网约车实时等待控制问题,提高匹配效率与乘车体验。
📝 Abstract
Real-time hold control is a high-leverage mechanism in large-scale ride-hailing systems: by selectively deferring driver-order pairs, the platform can wait for better matching opportunities and improve end-to-end passenger-driver experience. Existing production systems such as EXHOLD learn bandit-based hold policies from handcrafted combinations of trip completion, cancellations, waiting time, and driver effort. However, designing such rewards becomes increasingly difficult as marketplace preferences are heterogeneous and observed passenger-driver behavior can be sparse, noisy, and affected by dynamic supply-demand conditions. We present ALIGN-HOLD, a production-scale experience alignment framework that learns hold policy from implicit marketplace preferences. ALIGN-HOLD constructs complementary preference pairs from order trajectories, driver trajectories, and contemporaneous local matching graphs, and trains an experience Reward Model (RM) using balanced multi-view sampling and model-adaptive hard preference sampling. During simulator-based policy learning, the frozen RM provides a dense, context-dependent reward and supports label-free filtering of low-identifiability interactions whose behavioral feedback is difficult to attribute to matching quality. We deploy ALIGN-HOLD on DiDi's ride-hailing platform and evaluate it in a 28-day randomized A/B experiment, covering approximately 100,000 passenger requests per day. Compared with the deployed production policy, ALIGN-HOLD achieves statistically significant improvements in trip completion rate and driver income, while significantly reducing passenger cancellations before and after driver acceptance. Complementary ablations, RM diagnostics, and behavioral analyses validate the contributions of the proposed components. ALIGN-HOLD has been fully ramped up and is currently serving DiDi's Brazil marketplace.
Problem

Research questions and friction points this paper is trying to address.

Real-time hold control
Large-scale ride-hailing systems
Marketplace preferences
Heterogeneous preferences
Sparse and noisy behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

Experience Alignment
Real-time Hold Control
Reward Model
Multi-view Sampling
Hard Preference Sampling
🔎 Similar Papers
No similar papers found.
Z
Zuhao Zhang
Shanghai Jiao Tong University
X
Xu Liu
Didichuxing Co. Ltd
K
Kai Wan
Didichuxing Co. Ltd
Zihao Lu
Zihao Lu
Master Student, at University of Würzburg
Computer VisionImage ProcessingComputational PhotographyArtificial Intelligence
L
Li Ma
Didichuxing Co. Ltd
S
Shuai Li
Shanghai Jiao Tong University