Can We Trust Video Hallucination Detectors? VidHalLoc for Evaluating the Evaluators

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入VidHalLoc基准,使用2000个对抗性幻觉样本评估视频幻觉检测方法的可靠性,揭示现有检测器准确率仅34.63%。
📝 Abstract
Video-language models and video agents can produce hallucinations that conflict with spatiotemporal evidence. Existing benchmarks mainly evaluate model hallucinations, and heterogeneous mechanisms make detector reliability difficult to compare. We introduce VidHalLoc, a benchmark that evaluates hallucination detection methods under a unified diagnostic evaluation protocol using 2,000 adversarial hallucination samples across Video Question Answering and Video Captioning tasks, spanning Ontology and Dynamic hallucination categories. To construct VidHalLoc efficiently, we introduce VideoHALO, a Harness Engineering-informed multi-agent workflow that decomposes data construction into four executable stages supported by a memory system and a communication protocol. Evaluation of fifteen methods reveals that the four dedicated detectors peak at an Overall accuracy of only 34.63%, indicating limited reliability across video hallucination types [Dataset Repository: https://huggingface.co/datasets/wesfggfd/VidHalLoc].
Problem

Research questions and friction points this paper is trying to address.

Video Hallucination
Detection Methods
Reliability
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

VidHalLoc
VideoHALO
Hallucination Detection
Unified Diagnostic Evaluation Protocol
Adversarial Hallucination Samples
🔎 Similar Papers
No similar papers found.