Semantic-Spatial Agreement Verification for Mitigating Object Hallucination in Multimodal Large Language Models
为减少多模态大语言模型中的物体幻觉问题,提出了一种无需训练的语义-空间一致性验证方法SSAV,通过聚合多个提示和查询诱导区域验证来提高物体声明的准确性。
为减少多模态大语言模型中的物体幻觉问题,提出了一种无需训练的语义-空间一致性验证方法SSAV,通过聚合多个提示和查询诱导区域验证来提高物体声明的准确性。
本文研究了目标检测与参数估计的联合限制问题,通过引入后验熵体积和联合互信息两个概念,构建了一个统一的感知理论基础。
为解决异构设备联邦学习中模型配置受限问题,提出FANS框架及FPS算法,通过共享架构空间和并行训练优化模型性能。
This study addresses the challenge of attribute misbinding in large vision-language models within dense homogeneous scenes, where existing metrics prove inadequate. We formally define the DSCAM task and construct InstaBind-Lite, a controlled benchmark accompanied by a specialized evaluation framework. Through fine-grained instance annotation and multi-level question-answering design, this work enables quantitative assessment of attribute transfer while exposing critical blind spots in traditional evaluations. Experiments reveal misbinding rates of 19.84% for open-source models and 7.55% for API-based models, precisely localizing error sources. Ultimately, this research establishes a novel, traceable evaluation paradigm for assessing fine-grained attribute binding capabilities in large multimodal models.
为减少多模态大语言模型中的物体幻觉问题,提出了一种无需训练的语义-空间一致性验证方法SSAV,通过聚合多个提示和查询诱导区域验证来提高物体声明的准确性。
本文研究了目标检测与参数估计的联合限制问题,通过引入后验熵体积和联合互信息两个概念,构建了一个统一的感知理论基础。
为解决异构设备联邦学习中模型配置受限问题,提出FANS框架及FPS算法,通过共享架构空间和并行训练优化模型性能。
This study addresses the challenge of attribute misbinding in large vision-language models within dense homogeneous scenes, where existing metrics prove inadequate. We formally define the DSCAM task and construct InstaBind-Lite, a controlled benchmark accompanied by a specialized evaluation framework. Through fine-grained instance annotation and multi-level question-answering design, this work enables quantitative assessment of attribute transfer while exposing critical blind spots in traditional evaluations. Experiments reveal misbinding rates of 19.84% for open-source models and 7.55% for API-based models, precisely localizing error sources. Ultimately, this research establishes a novel, traceable evaluation paradigm for assessing fine-grained attribute binding capabilities in large multimodal models.