Adaptive Double-Booking Strategy for Outpatient Scheduling Using Multi-Objective Reinforcement Learning
This study addresses the significant inefficiencies and inequities in outpatient care caused by patient no-shows, which conventional fixed double-booking strategies fail to mitigate due to their inability to adapt to dynamic environments and individual heterogeneity. To overcome these limitations, we propose an adaptive double-booking framework grounded in multi-objective reinforcement learning. The approach integrates a multi-head attention-based soft random forest to predict individual no-show risk, which is then embedded into the state representation of a Markov decision process. We further design a multi-policy proximal policy optimization algorithm augmented with a KL divergence–based τ-rule to enable selective knowledge transfer across policies, enhancing both convergence and solution diversity. Additionally, SHAP values are employed to improve the interpretability of scheduling decisions. Experimental results demonstrate that our method substantially outperforms traditional heuristic strategies in alleviating clinic congestion and mitigating the adverse effects of no-shows.