Temporal-Causal Inference for Reinforcement Learning via Automata Learning

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入TCIRL框架解决强化学习中由隐藏时间模式控制的不可逆阶段转换问题,该方法联合学习控制策略并推断隐藏的时间因果关系。
📝 Abstract
We consider reinforcement learning in environments with dynamics that undergo an irreversible phase transition governed by a hidden temporal pattern. The agent observes the base state but cannot observe the phase directly. We formalize this problem as a two-phase non-Markovian decision process and introduce Temporal-Causal Inference for Reinforcement Learning (TCIRL), a framework that jointly learns a control policy and infers the hidden temporal cause of the phase transition. TCIRL maintains a hypothesis deterministic finite automaton (DFA) to track what phase is active and refines it via counterexample-driven SAT-based synthesis. We prove that the hypothesis converges almost surely to a DFA recognizing the true cause language on all attainable label sequences, yielding an optimal policy for the original non-Markovian decision process. Experiments on a genetic therapy gridworld and a traffic signal environment show that TCIRL recovers the correct cause DFA and matches the full-information baseline in both domains.
Problem

Research questions and friction points this paper is trying to address.

Temporal-Causal Inference
Reinforcement Learning
Non-Markovian Decision Process
Innovation

Methods, ideas, or system contributions that make the work stand out.

Temporal-Causal Inference
Reinforcement Learning
Deterministic Finite Automaton
Phase Transition
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jan Corazza
Research Center Trustworthy Data Science and Security, TU Dortmund University, Dortmund, 44227, Germany
D
Daniil Kaminskyi
Research Center Trustworthy Data Science and Security, TU Dortmund University, Dortmund, 44227, Germany
S
Simon Lutz
Research Center Trustworthy Data Science and Security, TU Dortmund University, Dortmund, 44227, Germany
P
Patrick Nossol
Research Center Trustworthy Data Science and Security, TU Dortmund University, Dortmund, 44227, Germany
H
Hadi Partovi Aria
School for Engineering of Matter, Transport, and Energy at Arizona State University, Tempe, AZ 85281, USA
Zhe Xu
Zhe Xu
Assistant Professor, Arizona State University
Cyber-Physical SystemsControl TheoryReinforcement LearningFormal MethodsRobotics
Daniel Neider
Daniel Neider
TU Dortmund University and Center for Trustworthy Data Science and Security
Formal MethodsMachine LearningLogicArtificial Intelligence