The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice

๐Ÿ“… 2026-08-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
็ ”็ฉถ้€š่ฟ‡็ฎ€ๅ•ๆŽข้’ˆๆ–นๆณ•ๆœ‰ๆ•ˆๆฃ€ๆต‹ๅคง่ง„ๆจก่ฏญ่จ€ๆจกๅž‹็š„ๅนป่ง‰้—ฎ้ข˜๏ผŒๅ‘็Žฐไฟกๅทไธป่ฆ็”ฑๅ•ไธ€ๅ‡ๅ€ผๅ็งป็ป„ๆˆ๏ผŒ็ฎ€ๅŒ–ไบ†ๅคๆ‚ๆžถๆž„็š„้œ€ๆฑ‚ใ€‚
๐Ÿ“ Abstract
Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across three 7B-scale models and three datasets in a paired-example paradigm, we find the signal overwhelmingly dominated by a single mean-shift component, and removing this direction collapses detection to chance. Shrinkage linear discriminant analysis closes about 73% of the gap between 1D and full-dimensional classifiers, so apparent architectural complexity largely reflects high-dimensional covariance estimation difficulty rather than exploitable non-linearity. A simple L2-regularized logistic regression (0.952 AUROC) bounds or outperforms twelve controlled architectural alternatives, and our multi-layer aggregation exceeds CLAP cross-layer attention probing under matched paradigm. Because the signal spans a contiguous layer band, LayerMix aggregates it to match oracle-layer performance without oracle access. Our claims characterize the geometry within the controlled paired-example paradigm. Our code is available at https://github.com/js-lee-AI/LayerMix.
Problem

Research questions and friction points this paper is trying to address.

Hallucination
Hidden-state Probes
Mean-shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

mean-shift
L2-regularized logistic regression
LayerMix
๐Ÿ”Ž Similar Papers
No similar papers found.