Adaptive Triggering for Bias Correction in LLM Reasoning

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出了一种基于在线变点检测的方法,通过自适应触发机制在大语言模型推理过程中纠正偏见,有效平衡了干预时机与准确性。
📝 Abstract
Chain-of-thought prompting can expose and amplify demographic stereotypes within an LLM's intermediate reasoning and create a failure mode that final-answer debiasing alone cannot address. Mitigating such bias during generation presents a fundamental timing problem: intervening too late allows biased reasoning to propagate, while unnecessarily intervening can disrupt otherwise correct reasoning. Existing approaches largely avoid this decision by either evaluating completed reasoning chains post hoc or intervening at predetermined steps, leaving open when a developing reasoning trajectory provides sufficient evidence to warrant correction. We formulate this decision as an online change-point detection problem. A per-step bias signal updates a CUSUM statistic and a targeted correction is injected only when accumulated evidence crosses a detector-specific threshold calibrated on held-out data. We instantiate the framework with a white-box signal derived from next-token probabilities and a black-box signal obtained from an LLM judge, enabling deployment with both open-weight and hosted models. On gpt-4o-mini adaptive black-box triggering recovers most of the disambiguated-context accuracy lost under fixed-interval intervention while requiring substantially fewer interventions. That result holds even with an independent judge. Across six open-weight models, the white-box signal improves ambiguous-item accuracy on all six but reduces disambiguated-item accuracy on five because it cannot distinguish unsupported stereotype reliance from correct, stereotype-congruent evidence.
Problem

Research questions and friction points this paper is trying to address.

Bias Correction
Chain-of-thought Prompting
Large Language Models
Online Change-point Detection
Intermediate Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Triggering
Bias Correction
Online Change-Point Detection
CUSUM Statistic
Chain-of-Thought Prompting
🔎 Similar Papers
No similar papers found.
N
Nayoung Kim
School of Computing and Augmented Intelligence, Arizona State University, Tempe, AZ, USA
M
Mickey Mancenido
School of Mathematical and Natural Sciences, Arizona State University, Phoenix, AZ, USA
Huan Liu
Huan Liu
Regents Professor, Arizona State University
AIData MiningFeature SelectionSocial ComputingSocial Media Mining