From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究利用语义熵和令牌不确定性两种信号检测黑盒LLM中的幻觉问题,并提出TopK、CoCoA、Gated和Stacked四种方法来提高检测准确性。
📝 Abstract
When LLMs support public-facing or high-stakes workflows, missed fabrications can harm users and institutions, while false alarms consume limited human-review capacity. When no trusted context or reference document is available, we study two signals accessible through black-box model APIs: semantic entropy, which measures disagreement among sampled response meanings, and uncertainty derived from token log-probabilities. Their failure modes can be complementary: semantic entropy becomes uninformative when responses form one semantic cluster, while token uncertainty can miss consistently confident errors. We extend token-based uncertainty detection by aggregating token-level signals across sampled responses through our TopK method, evaluate the hybrid CoCoA method, which combines target-response uncertainty with semantic dissimilarity, and propose and study two supervised methods: Gated, which routes single-cluster cases to an aggregated-token-feature classifier, and Stacked, which learns jointly from semantic uncertainty and broader token features. We evaluate seven benchmarks, including five public benchmarks (four text datasets and multimodal handwritten-cheque extraction) and two constructed benchmarks (Financial Summaries and Long-Text QA), using four language models. In our evaluation across models and datasets, Stacked gave the best performance in nearly half of the cases, while TopK and CoCoA remain competitive without supervised training labels, although their thresholds require careful calibration. No method is universally strongest. We therefore evaluate performance at false-positive-rate budgets from 1% to 15%, assess their sensitivity to generation and calibration choices, and examine variation across dataset characteristics.
Problem

Research questions and friction points this paper is trying to address.

Hallucination Detection
Black-Box LLMs
Semantic Entropy
Token Uncertainty
False Alarms
Innovation

Methods, ideas, or system contributions that make the work stand out.

semantic entropy
token log-probabilities
TopK method
CoCoA method
Stacked method
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Urja Pawar
Urja Pawar
PhD Student (2019-2024), MTU
AIMLXAIHealthcare
R
Rajitha Ramanayake
BNY
O
Owen O'Neill
BNY
N
Nabeel Kemal
BNY
A
Abhishek Mandal
BNY
H
Houssem Chatbri
BNY
C
Christopher Martin
BNY