Focus Where It Counts: A Salience-Driven Vision-Language Model for Low Vision Assistance
为解决视障辅助技术中信息优先级问题,本文提出一种基于显著性的视觉-语言模型Salience-LLaVA,并构建了相关数据集以优化场景描述。
为解决视障辅助技术中信息优先级问题,本文提出一种基于显著性的视觉-语言模型Salience-LLaVA,并构建了相关数据集以优化场景描述。
This paper addresses key challenges in multi-label chest X-ray classification—extreme label imbalance, asymmetric misdiagnosis costs, and inadequate modeling of label co-occurrence—by proposing a lightweight, efficient diagnostic framework. The core innovation is a learnable sparse label graph refinement module that explicitly captures label dependencies at the logits level via single-step message passing, introducing negligible computational overhead. The method integrates a SE-ResNeXt101 backbone with asymmetric loss, mixed-precision training, cosine annealing, and exponential moving average (EMA) for robust optimization, and employs multi-fold cross-validation with test-time augmentation (TTA) ensembling. Experiments demonstrate significant improvement in macro-AUC, achieving high performance without additional annotations. The framework is computationally efficient, hardware-friendly, and clinically deployable.
为解决视障辅助技术中信息优先级问题,本文提出一种基于显著性的视觉-语言模型Salience-LLaVA,并构建了相关数据集以优化场景描述。
This paper addresses key challenges in multi-label chest X-ray classification—extreme label imbalance, asymmetric misdiagnosis costs, and inadequate modeling of label co-occurrence—by proposing a lightweight, efficient diagnostic framework. The core innovation is a learnable sparse label graph refinement module that explicitly captures label dependencies at the logits level via single-step message passing, introducing negligible computational overhead. The method integrates a SE-ResNeXt101 backbone with asymmetric loss, mixed-precision training, cosine annealing, and exponential moving average (EMA) for robust optimization, and employs multi-fold cross-validation with test-time augmentation (TTA) ensembling. Experiments demonstrate significant improvement in macro-AUC, achieving high performance without additional annotations. The framework is computationally efficient, hardware-friendly, and clinically deployable.