Focus Where It Counts: A Salience-Driven Vision-Language Model for Low Vision Assistance

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视障辅助技术中信息优先级问题,本文提出一种基于显著性的视觉-语言模型Salience-LLaVA,并构建了相关数据集以优化场景描述。
📝 Abstract
Vision-language models (VLMs) are rapidly progressing and offer promising capabilities for assistive technologies supporting persons with blindness or low vision. However, existing VLMs are primarily designed for general-purpose captioning and do not explicitly model human perceptual priorities, thereby limiting their ability to emphasize the most relevant information in a scene. To address this gap, we propose a salience-driven captioning framework that prioritizes scene elements according to their importance for human-centered assistance. We curate three salience-aware datasets, namely, Salience COCO, Salience Flickr, and Salience VizWiz, with object-level salience annotations designed to reflect the visual information most relevant to low vision users across different environments. Building on these datasets, we introduce Salience-LLaVA, a salience-aware VLM that incorporates salience cues to generate captions in which important elements are mentioned in the order of importance. Our work makes four main contributions. We build salience-aware datasets verified by low vision participants, propose Salience-LLaVA to describe objects in the order of importance, introduce SCMI to evaluate ordering accuracy, and deploy the system on assistive glasses to demonstrate real-world practicality. Code and datasets are available at: https://github.com/topo-focus/Topofocus
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Low Vision Assistance
Salience
Perceptual Priorities
Captioning
Innovation

Methods, ideas, or system contributions that make the work stand out.

salience-driven
vision-language model
low vision assistance
Salience-LLaVA
SCMI
🔎 Similar Papers
No similar papers found.
J
Jiazhao Liang
New York University Tandon School of Engineering, Brooklyn, NY 11201, USA
Hao Huang
Hao Huang
New York University
Computer VisionArtificial IntelligenceRobotics
S
Shuaihang Yuan
NYUAD Center for Artificial Intelligence and Robotics (CAIR) and Embodied AI and Robotics (AIR) Lab, New York University Abu Dhabi, Abu Dhabi 129188, UAE
C
Congcong Wen
NYUAD Center for Artificial Intelligence and Robotics (CAIR) and Embodied AI and Robotics (AIR) Lab, New York University Abu Dhabi, Abu Dhabi 129188, UAE
G
Geeta Chandra Raju Bethala
NYUAD Center for Artificial Intelligence and Robotics (CAIR) and Embodied AI and Robotics (AIR) Lab, New York University Abu Dhabi, Abu Dhabi 129188, UAE
Giles Hamilton-Fletcher
Giles Hamilton-Fletcher
Research Scientist, NYU Langone Health
sensory substitutioncross-modal correspondencesqualiasynaesthesia
Yu Hao
Yu Hao
New York University
3D Vision
J
John-Ross Rizzo
NYU Grossman School of Medicine, NYU Langone Health, New York, NY 10016, USA
Mengyu Wang
Mengyu Wang
Assistant Professor, Harvard Medical School
Artificial IntelligenceMachine LearningOphthalmologyGlaucomaComputational Mechanics
Anthony Tzes
Anthony Tzes
Professor of Electrical and Computer Engineering, New York University Abu Dhabi, UAE
Control ApplicationsRobotics
Yi Fang
Yi Fang
Associate Professor of NYU Abu Dhabi and NYU Tandon
3D Computer Vision3D Deep Learning3D Meta Learning