Beyond Pure Sampling: Hybrid Optimization Mechanisms for Non-Convex Model Predictive Control
This work addresses the challenge of navigating complex cost landscapes in non-convex model predictive control, where nonlinear dynamics and multiple obstacles often trap gradient-based methods in suboptimal local minima. To overcome this limitation, we propose a Maximum Entropy Differential Dynamic Programming (ME-DDP) framework that integrates deterministic optimization with entropy-maximizing sampling. Our approach employs a two-stage mechanism: it first performs local gradient-based refinement via DDP and then leverages the inverse Hessian of the action-value function to guide policy sampling, enabling escape from local minima and balancing global exploration with local exploitation. We develop three ME-DDP variants, elucidate their theoretical connections to Model Predictive Path Integral (MPPI) control, and demonstrate superior performance across four navigation benchmarks—achieving higher success rates in high-dimensional systems, outperforming MPPI in low-dimensional settings, and exhibiting robustness in real-world quadrotor experiments through dense obstacle fields.