The Self-Correction Illusion: LLMs Correct Others but Not Themselves

📅 2026-06-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates why large language models struggle to correct their own reasoning errors yet effectively fix identical errors when attributed to external roles. By fixing erroneous content and systematically varying only the role labels (e.g., <thought>, user, tool, <memory>) in dialogue templates, the authors isolate the effect of role attribution. Experiments across 13 model–domain combinations employ SHA-256 checksums to ensure content consistency and include cross-model, cross-task comparisons with statistical significance testing. Results show that relabeling errors from <thought> to an external role increases explicit correction rates by 23–93 percentage points (10 comparisons, p<0.001), with <memory> most effective in mathematical tasks and user in logical reasoning. The findings indicate that self-correction failure stems from representational bias induced by role labels—not inherent cognitive limitations—and propose a training-free, prompt-structure intervention to mitigate this issue.
📝 Abstract
Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical claims appear under external sources. We ask whether this asymmetry reflects a capability deficit or a role-label artifact: does an agent's willingness to correct a wrong claim depend causally on the chat-template role that carries it, rather than on the claim's content? Our setup keeps the erroneous claim byte-identical across all conditions (SHA-256 verified) and varies only its wrapping role: the agent's own \role{<thought>}, a \role{user} message, a \role{tool} response, or a \role{system <memory>} block. Across 13 model-domain cells covering seven model families and three domains ($n{=}30$ paired tasks per cell), relabeling the claim from \role{<thought>} to an external role lifts the explicit-correction rate by 23 to 93 percentage points, with 10 of 13 cells reaching $p{<}0.001$. Further experiments confirm that the effect is asymmetric, mechanistically decomposable, and robust across domains. The failure to self-correct is not a cognitive deficit; it is a chat-template artifact. We exploit this artifact by designing a prompt-structure-only intervention that requires no training and no model modification, with its strongest role label being domain-dependent: \role{<memory>} dominates on math, while a plain \role{user} message dominates on logical deduction.
Problem

Research questions and friction points this paper is trying to address.

self-correction
large language models
role labeling
reasoning errors
chat template
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-correction illusion
role-label artifact
chat-template
prompt engineering
large language models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kuan-Yen Chen
National Cheng Kung University
Fang-Yi Su
Fang-Yi Su
National Cheng Kung University, PhD student
AI for MedicineGenerative ModelArtificial Intelligence
J
Jung-Hsien Chiang
National Cheng Kung University