OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
OmniHallu框架通过多代理架构和模态特定专家验证,解决了多模态大语言模型在跨模态理解和生成任务中的幻觉检测问题。
📝 Abstract
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse tasks, they suffer from hallucinations where generated outputs contradict or misrepresent input semantics. Existing research typically addresses hallucination detection within a single modality or task type, limiting generalizability. We introduce OmniHallu, a unified hallucination detection framework spanning both comprehension and generation tasks across image, video, and audio modalities. We contribute OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations covering six cross-modal tasks: image-to-text (I2T), video-to-text (V2T), audio-to-text (A2T), text-to-image (T2I), text-to-video (T2V), and text-to-audio (T2A). Our multi-agent architecture decomposes model outputs into atomic claims, verifies them through modality-specific experts, and aggregates evidence via structured reasoning. We further propose a preference-optimized trainable verifier that approximates the multi-agent decision boundary, reducing expert calls by 66% with minimal performance loss. Extensive experiments reveal a consistent modality-dependent performance gradient and provide fine-grained insights into cross-modal hallucination patterns.
Problem

Research questions and friction points this paper is trying to address.

hallucination detection
multimodal large language models
cross-modal tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Hallucination Detection
Cross-Modal Comprehension and Generation
Multi-Agent Architecture
Preference-Optimized Verifier
OmniHallu-Bench
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Jianjiang Yang
Jianjiang Yang
Master of Science in Computer Science, University of Bristol
MLLMs
P
Peihang Li
The University of Hong Kong
S
Shanqing Xu
Huazhong University of Science and Technology
M
Mengchen Qian
Huazhong University of Science and Technology
L
Lu Zhang
Shanghai Academy of Educational Sciences
Meng Luo
Meng Luo
National University of Singapore
Human-Centered AIMultimodal UnderstandingMultimodal Reasoning