Rewarding Curse: Analyze and Mitigate Reward Modeling Issues for LLM Reasoning

📅 2025-03-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) exhibit inconsistency between effectiveness and faithfulness in chain-of-thought (CoT) reasoning, primarily due to biases in reward modeling. Method: We propose the first holistic analytical framework jointly modeling the “question–CoT–answer” information interaction, leveraging information quantification, faithfulness diagnosis, and information gain assessment to uncover how question difficulty, information gain, and directional information flow critically influence CoT performance. We identify a prevalent failure mode wherein models bypass incomplete CoTs and directly retrieve answers from the question—yielding superficially correct but logically unfaithful outputs. Building on these insights, we design a novel CoT generation and evaluation algorithm driven by question-informed information backtracking and information gain optimization. Contribution/Results: Evaluated across diverse reasoning tasks, our approach significantly improves CoT faithfulness (+21.3%) and effectiveness (+14.7% accuracy), empirically validating the centrality of information backtracking in faithful reasoning.

Technology Category

Application Category

📝 Abstract
Chain-of-thought (CoT) prompting demonstrates varying performance under different reasoning tasks. Previous work attempts to evaluate it but falls short in providing an in-depth analysis of patterns that influence the CoT. In this paper, we study the CoT performance from the perspective of effectiveness and faithfulness. For the former, we identify key factors that influence CoT effectiveness on performance improvement, including problem difficulty, information gain, and information flow. For the latter, we interpret the unfaithful CoT issue by conducting a joint analysis of the information interaction among the question, CoT, and answer. The result demonstrates that, when the LLM predicts answers, it can recall correct information missing in the CoT from the question, leading to the problem. Finally, we propose a novel algorithm to mitigate this issue, in which we recall extra information from the question to enhance the CoT generation and evaluate CoTs based on their information gain. Extensive experiments demonstrate that our approach enhances both the faithfulness and effectiveness of CoT.
Problem

Research questions and friction points this paper is trying to address.

Analyze factors affecting Chain-of-Thought (CoT) effectiveness in reasoning tasks.
Identify and address unfaithful CoT issues in LLM reasoning processes.
Propose algorithm to enhance CoT faithfulness and effectiveness using information recall.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Analyzes CoT effectiveness via problem difficulty, information gain, flow
Identifies unfaithful CoT through question-CoT-answer interaction analysis
Proposes algorithm enhancing CoT by recalling extra question information
🔎 Similar Papers
No similar papers found.
J
Jiachun Li
School of Artificial Intelligence, University of Chinese Academy of Sciences; The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences
P
Pengfei Cao
School of Artificial Intelligence, University of Chinese Academy of Sciences; The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences
Yubo Chen
Yubo Chen
Institute of Automation, Chinese Academy of Sciences
Natural Language ProcessingInformation ExtractionEvent ExtractionLarge Language Model
J
Jiexin Xu
China Merchants Bank
H
Huaijun Li
China Merchants Bank
X
Xiaojian Jiang
China Merchants Bank
K
Kang Liu
School of Artificial Intelligence, University of Chinese Academy of Sciences; The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences
J
Jun Zhao
School of Artificial Intelligence, University of Chinese Academy of Sciences; The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences