Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文解决了大型视觉-语言模型中的幻觉检测问题,通过构建一个模型无关的检测任务,并使用五类幻觉分类法在四种语言中进行评估。
📝 Abstract
In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration \textbf{M}istakes in \textbf{Vision} language model\textbf{s}), which is hosted at the UncertaiNLP Workshop co-located with EMNLP 2026. Following the success of the 2024 and 2025 tasks, this time we aim to tackle hallucinations through a model-agnostic detection task focused on large vision-language models. Building on the recently introduced SHEEP dataset, designed for long-term evaluation across model generations, the task invites participants to detect and classify fine-grained hallucination spans in image-conditioned text generation (VQA, image captioning, etc.). The evaluation uses a five-class taxonomy of hallucinations spanning four languages: Chinese, English, French, and Italian. The shared task generated strong interest in the NLP community worldwide, with 27 teams contributing 600+ system submissions. The best systems achieve average scores of 0.58 in character-level correlation, 0.46 in label-conditioned correlation, and 0.51 in intersection-over-union (IoU) across four languages, outperforming the baselines by 30-40 points.
Problem

Research questions and friction points this paper is trying to address.

hallucination detection
vision-language models
multilingual
fine-grained classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

hallucination detection
large vision-language models
SHEEP dataset
multilingual
🔎 Similar Papers