VisER: Visual Evidence and Reliance for Object Hallucination Detection in LVLMs

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大视觉-语言模型中的对象幻觉问题,提出VisER方法,通过视觉证据和依赖性两个角度评估生成对象的真实性。
📝 Abstract
Object hallucination remains a persistent reliability issue in large vision-language models, where generated object mentions may sound plausible but lack visual grounding. Recent training-free detectors use internal signals such as token likelihood, attention, visual confidence, or image-text similarity to identify hallucinated objects. These signals are useful, but they are often source-confounded. They measure how strongly an object is supported inside the model without distinguishing whether that support comes from object-specific visual evidence or the generated text prefix. In difficult cases, a hallucinated object can still receive high internal support because it fits the scene, is associated with nearby visual cues, or follows naturally from the generated text prefix. We propose VisER, a training-free two-sided metric for object-level hallucination detection. VisER evaluates each generated object mention from two complementary views. Visual Evidence measures whether object-context compatibility is backed by object-specific evidence from image tokens. Visual Reliance measures whether the object is supported more by the image than by the generated prefix. Combining these views gives a more source-aware grounding score, while avoiding additional object-level verification generations. Across multiple LVLMs and benchmarks, VisER improves AUROC and AUPR over a range of baselines.
Problem

Research questions and friction points this paper is trying to address.

object hallucination
large vision-language models
visual grounding
internal signals
source-confounded
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Evidence
Visual Reliance
Training-free Detector
Object Hallucination Detection
LVLMs
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Afsaneh Hasanebrahimi
School of Computing and Information Systems, The University of Melbourne, Melbourne, Australia
Hanxun Huang
Hanxun Huang
The University of Melbourne
Trustworthy AIAI SafetyGenerative AICyber Security
Christopher Leckie
Christopher Leckie
Professor, Computing and Information Systems, The University of Melbourne
artificial intelligencemachine learninganomaly detectionclusteringcyber security
S
Sarah Erfani
School of Computing and Information Systems, The University of Melbourne, Melbourne, Australia