🤖 AI Summary
本文通过使用多个微调的视觉-语言模型作为独立注释器,并结合字符级多数投票和激活探针方法,解决了视觉-语言模型输出中幻觉字符跨度检测与分类的问题。
📝 Abstract
This paper describes our submission to the SHROOM-Visions shared task on detecting and classifying hallucinated character spans in vision-language model outputs across four languages. We employ several fine-tuned vision-language models as independent annotators and combine their span predictions through character-level majority voting, and additionally explore activation probes. The approach ranks first in three of four languages and places on the podium in every language and metric. Our analysis indicates that disagreement among diverse models tracks disagreement among human annotators.