NeuroActiSep: Detecting Factual Hallucinations from Feed-Forward Neurons in a Single Pass

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种通过单次检测前馈神经元来识别大语言模型中事实性幻觉的方法,使用定制数据集选择并转移相关神经元以训练分类器。
📝 Abstract
Hallucination in large language models reduces their reliability and slows adoption. Various white-box studies have used internal representations to detect patterns of truthfulness and factuality. A less-studied approach is to identify feed-forward neurons correlated with hallucination. We propose a method to rank feed-forward neurons at the final prompt token using a custom neuron selection dataset. We transfer the selected neuron identities to train hallucination classifiers on other factual question answering datasets. Our work provides empirical evidence that probes trained using the features from the selected neurons perform on par with probes trained on internal states. We also analyze the distribution of selected neurons and the effect of layer depth on detection performance.
Problem

Research questions and friction points this paper is trying to address.

Hallucination
Large Language Models
Factuality
Innovation

Methods, ideas, or system contributions that make the work stand out.

feed-forward neurons
hallucination detection
single pass
neuron ranking
factual question answering
🔎 Similar Papers
No similar papers found.
A
Ali Derogar Odolou
School of Intelligent Systems Engineering, College of Interdisciplinary Science and Technology, University of Tehran, Tehran, Iran
R
Reza Nazari
School of Intelligent Systems Engineering, College of Interdisciplinary Science and Technology, University of Tehran, Tehran, Iran
Mostafa Salehi
Mostafa Salehi
Associate Professor, University of Tehran
Social Network and Media AnalysisNetwork Science