Beyond Bellman: High-Order Generator Regression for Continuous-Time Policy Evaluation

📅 2026-04-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses policy evaluation in continuous time over a finite horizon using discrete trajectories, where conventional Bellman-based methods suffer from only first-order accuracy. The authors propose a high-order generator regression framework that eliminates low-order truncation errors by matching moments of multi-step transitions and integrates backward parabolic equation regression to construct a high-order policy evaluator. This approach provides the first systematic elimination of truncation errors in continuous-time policy evaluation and establishes a theoretical characterization of error decomposition and gain identifiability. Empirical results demonstrate that the proposed second-order estimator significantly outperforms Bellman baselines across multiple benchmark tasks, with notably improved evaluation accuracy within the theoretically predicted effective regime.

Technology Category

Application Category

📝 Abstract
We study finite-horizon continuous-time policy evaluation from discrete closed-loop trajectories under time-inhomogeneous dynamics. The target value surface solves a backward parabolic equation, but the Bellman baseline obtained from one-step recursion is only first-order in the grid width. We estimate the time-dependent generator from multi-step transitions using moment-matching coefficients that cancel lower-order truncation terms, and combine the resulting surrogate with backward regression. The main theory gives an end-to-end decomposition into generator misspecification, projection error, pooling bias, finite-sample error, and start-up error, together with a decision-frequency regime map explaining when higher-order gains should be visible. Across calibration studies, four-scale benchmarks, feature and start-up ablations, and gain-mismatch stress tests, the second-order estimator consistently improves on the Bellman baseline and remains stable in the regime where the theory predicts visible gains. These results position high-order generator regression as an interpretable continuous-time policy-evaluation method with a clear operating region.
Problem

Research questions and friction points this paper is trying to address.

continuous-time policy evaluation
Bellman equation
high-order regression
time-inhomogeneous dynamics
discrete trajectories
Innovation

Methods, ideas, or system contributions that make the work stand out.

high-order generator regression
continuous-time policy evaluation
moment-matching
backward parabolic equation
Bellman recursion
🔎 Similar Papers
No similar papers found.
Yaowei Zheng
Yaowei Zheng
Ph.D. student, Beihang University
Machine LearningNatural Language Processing
Richong Zhang
Richong Zhang
Professor of Computer Science, Beihang University
Data MiningRecommender SystemSocial Computing
S
Shenxi Wu
Fudan University, China
S
Shirui Bian
Fudan University, China
H
Haosong Zhang
Fudan University, China
Li Zeng
Li Zeng
Peking University
LLM training and inferenceVector ComputingGraph Computing
X
Xingjian Ma
Fudan University, China
Y
Yichi Zhang
Stern School of Business, New York University, USA