🤖 AI Summary
This work addresses the challenge of balancing information freshness and correctness in pull-based remote state estimation by formulating a Markov decision process with discounted cost, using the Age of Incorrect Information (AoII) as the optimization objective. By revealing that the belief state depends only on the most recent successful observation and the subsequent duration without updates, the authors propose belief compression, truncation-based approximation, and a hybrid estimator, along with an early steady-switching mechanism. The framework is extended to multi-source settings, where indexability conditions are established and a Whittle index–based heuristic policy is derived. Theoretical and empirical results demonstrate that the optimal single-source policy exhibits a lookup-table waiting structure, while the proposed multi-source scheduling heuristic achieves near-optimal performance with significantly reduced computational overhead and provides computable performance bounds.
📝 Abstract
We study pull-based remote state estimation of an arbitrary, multi-state Markov source while accounting for both freshness and correctness attributes of information. To that end, we formulate a discounted optimization problem in terms of the age of incorrect information (AoII), and express it as a joint source-AoII belief Markov decision process (MDP) under maximum a posteriori (MAP) estimation. We then exploit the information structure of the model and prove that every reachable belief is represented by the last successfully observed source state and the number of time slots elapsed since that observation. For numerical computation, we truncate the elapsed no-success duration at a finite level and derive an explicit error bound and a criterion for selecting the truncation parameter. For reliable links, we show that an optimal policy can be represented by a look-up table of waiting times. For unreliable links, we propose a persistent policy and derive computable performance bounds. We also show that the MAP estimate stabilizes after a finite number of time slots. To further reduce memory requirements, we introduce a hybrid estimator with an early stationary switch and derive a computable bound on the resulting difference in performance. Finally, we extend the framework to multiple sources, formulate the scheduling problem as a restless multi-armed bandit, establish a sufficient condition for indexability, and develop an approximate Whittle index policy based on interpolation. Our numerical results illustrate the structure of the optimal single-source policy, evaluate the performance of the multi-source policies, and verify that the proposed heuristic policies closely approach the optimal solution while substantially reducing computational efforts.