Target Discounted Sum Problem on Markov Chains with Applications to Markov Decision Processes

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了马尔可夫链上的目标折扣和问题,通过证明特定路径事件概率为零,并应用自动机理论技术,进而应用于马尔可夫决策过程。
📝 Abstract
The discounted sum is a way to aggregate a sequence of weights from a finite alphabet $Σ$, i.e., for a discount factor $λ$, the discounted sum of a sequence $w_0 w_1 w_2 \cdots$ over $Σ$ is $\sum_{i \in \mathbb{N}} w_i λ^i$. The target discounted-sum problem, which is currently open, asks, given $λ,Σ$ and a target $t$, whether there exists an infinite sequence over $Σ$ whose discounted sum is equal to $t$. We study and solve a probabilistic variant of this problem, i.e., the target discounted-sum problem on Markov chains. To do this, we prove that the event consisting of paths whose discounted sum is equal to the target and has infinitely many distinct suffix sums has probability zero. This structural property allows us to solve the target discounted-sum problem on Markov chains using an automata-theoretic technique. We apply our technical results to Markov decision processes with target discounted-sum objectives: we show that the infimum value and the finite-memory supremum value are computable in pseudo-polynomial time and are attained by deterministic finite-memory strategies.
Problem

Research questions and friction points this paper is trying to address.

discounted sum
Markov chains
target discounted-sum problem
Innovation

Methods, ideas, or system contributions that make the work stand out.

target discounted-sum problem
Markov chains
automata-theoretic technique
Markov decision processes
deterministic finite-memory strategies