Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
Existing methods struggle to precisely identify the internal mechanisms within large language models that support complex reasoning and fail to effectively model the sequential influence from internal components to final outputs. This work proposes the Integrated Policy Gradient (IPG) framework, which introduces— for the first time—the policy gradient concept from reinforcement learning into the interpretability research of large language models. By backpropagating composite signals such as reasoning outcomes and incorporating retrospective analysis of reasoning trajectories, IPG identifies and modulates neurons or modules that cumulatively contribute to long-range reasoning. Experiments demonstrate that IPG achieves more accurate mechanistic localization across multiple reasoning models and effectively tunes both the capability and intensity of reasoning, thereby validating its efficacy and generalizability.