An MDP Model for Censoring in Harvesting Sensors: Optimal and Approximated Solutions
This paper addresses energy-efficient information transmission for energy-harvesting sensors, aiming to maximize cumulative message utility (i.e., importance) under finite energy constraints. Method: We formulate the problem as an infinite-horizon Markov decision process (MDP) and—under a realistic battery dynamics model—rigorously prove that the optimal policy is a state-dependent importance-threshold truncation policy, where transmission decisions depend dynamically on the current battery level. Building on this structural insight, we propose a low-complexity, fast-converging model-driven stochastic approximation algorithm and benchmark it against Q-learning. Results: Experiments in both single-hop and multi-hop networks demonstrate that our algorithm significantly reduces computational overhead and accelerates convergence while achieving utility performance close to the theoretical optimum.