Institution profile

Shenyang Aerospace University

Academic institutionasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

Aug 17, 2026

This study addresses the challenge of attribute misbinding in large vision-language models within dense homogeneous scenes, where existing metrics prove inadequate. We formally define the DSCAM task and construct InstaBind-Lite, a controlled benchmark accompanied by a specialized evaluation framework. Through fine-grained instance annotation and multi-level question-answering design, this work enables quantitative assessment of attribute transfer while exposing critical blind spots in traditional evaluations. Experiments reveal misbinding rates of 19.84% for open-source models and 7.55% for API-based models, precisely localizing error sources. Ultimately, this research establishes a novel, traceable evaluation paradigm for assessing fine-grained attribute binding capabilities in large multimodal models.

0 citationsRead paper
Recent publications

Latest Papers

Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models

Aug 17, 2026

This study addresses the challenge of attribute misbinding in large vision-language models within dense homogeneous scenes, where existing metrics prove inadequate. We formally define the DSCAM task and construct InstaBind-Lite, a controlled benchmark accompanied by a specialized evaluation framework. Through fine-grained instance annotation and multi-level question-answering design, this work enables quantitative assessment of attribute transfer while exposing critical blind spots in traditional evaluations. Experiments reveal misbinding rates of 19.84% for open-source models and 7.55% for API-based models, precisely localizing error sources. Ultimately, this research establishes a novel, traceable evaluation paradigm for assessing fine-grained attribute binding capabilities in large multimodal models.

0 citationsRead paper