Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了结合大语言模型与强化学习时奖励信号理论地位不明确的问题,通过将LLM反馈作为有界势函数,确保即使在LLM评分不准确的情况下也能保持最优策略集。
📝 Abstract
Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward signals is often left implicit. We formalize the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and show that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term preserves the optimal policy set even when the LLM scores are inaccurate. This guarantee is stronger than what general LLM-as-reward approaches provide. We verify the result numerically on a small MDP under four potential configurations, including an adversarial one scaled to twenty times the base reward magnitude.
Problem

Research questions and friction points this paper is trying to address.

reward shaping
large language models
reinforcement learning
policy invariance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy-Invariant
Reward Shaping
Large Language Models
Hybrid RL Agents
Bounded Potential Function
C
Christophe D. Hounwanou
African Institute for Mathematical Sciences, Rwanda; AI Research and Innovation Nexus for Africa (AIRINA Labs), AI.Technipreneurs, Bénin
J
John Emeka Eze
African Institute for Mathematical Sciences, Rwanda
Y
Yaé U. Gaba
AI Research and Innovation Nexus for Africa (AIRINA Labs), AI.Technipreneurs, Bénin; Sefako Makgatho Health Sciences University (SMU), South Africa; African Center for Advanced Studies (ACAS), Cameroon