Algorithmic Optimality Guarantees for Nonsmooth $H_\infty$ Output-Feedback Policy Search

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究了非凸且非光滑的连续时间全阶动态输出反馈H∞策略搜索问题,通过扩展凸提升方法证明了ε-平稳性与O(ε)次优性的关系,并提出了一种非严格可行性二分法以获得ε最优稳定控制器。
📝 Abstract
We study continuous-time full-order dynamic output-feedback $H_\infty$ policy search, a nonconvex and nonsmooth problem. Direct policy search is a central paradigm in reinforcement learning and continuous control, but rigorous guarantees remain scarce in robust output-feedback settings. The $H_\infty$ problem is a canonical benchmark because it captures disturbance attenuation and robustness while exposing the hard nonsmooth geometry of policy-space optimization. We prove that on the exact identity-gauge slice of the extended convex lift, $\varepsilon$-stationarity yields $O(\varepsilon)$-suboptimality on compact exact slices, which in turn yields convergence-rate guarantees for nonsmooth policy-search methods. This result addresses the finite-time optimality-gap question raised by Guo and Hu [2022] in the more general dynamic output-feedback $H_\infty$ policy-search setting. We further use the established value equivalence supplied by extended convex lifting to formulate a nonstrict-feasibility bisection method with one final strict-feasibility recovery step, yielding an explicit $\varepsilon$-optimal stabilizing controller. These results provide a quantitative and algorithmic strengthening of prior qualitative optimality theory for nonsmooth $H_\infty$ policy search.
Problem

Research questions and friction points this paper is trying to address.

nonconvex
nonsmooth
H∞ control
output-feedback
policy search
Innovation

Methods, ideas, or system contributions that make the work stand out.

nonsmooth H∞ policy search
extended convex lifting
ε-stationarity
nonstrict-feasibility bisection method
🔎 Similar Papers
No similar papers found.