ViCo-SAM3: Vision-Conditioned Alignment for Open-Vocabulary Camouflaged Object Segmentation
为解决OVCOS中视觉与文本语义鸿沟问题,提出ViCo-SAM3框架,通过视觉条件模块动态调整文本嵌入,增强跨模态交互。
为解决OVCOS中视觉与文本语义鸿沟问题,提出ViCo-SAM3框架,通过视觉条件模块动态调整文本嵌入,增强跨模态交互。
为解决全景显著物体检测中的几何失真问题,提出TDFNet,通过三投影变形融合网络及交叉投影变形注意力模块提升检测性能。
该研究针对网络事件序列中的因果关系发现难题,提出了一种鲁棒的并发因果发现方法RCCD,通过引入注意力机制和自监督掩码重构框架来提高对缺失数据的鲁棒性和准确性。
This study addresses the prevalent issues of coarse, misaligned, or incomplete manual annotations in remote sensing semantic segmentation, which often distort model evaluation. To tackle this, the authors propose a training-free, reference-free mask fidelity assessment method that constructs counterfactual image pairs—preserving and erasing the region within the mask—and leverages a frozen vision-language model to evaluate whether class-specific evidence is concentrated inside the mask and absent outside it. This approach enables, for the first time, reference-free auditing of annotation quality in remote sensing segmentation, revealing systematic labeling biases across categories and facilitating automatic refinement of supervision signals. The proposed Contrastive Mask Fidelity (CMF) metric achieves 81% agreement with expert judgments across ten remote sensing datasets, substantially outperforming existing methods, and CMF-guided supervision significantly enhances cross-domain transfer performance.
This work addresses the challenge of effectively segmenting slender, anisotropic defects—such as cracks and scratches—on steel surfaces, which existing methods struggle to handle accurately. To this end, the authors propose the SPDCN network, which incorporates a Fuzzy-enhanced Multi-scale Context Module (FMCM) to adaptively fuse multi-scale contextual information. Additionally, an Adaptive Direction-Aware Deformable Convolution (ADADC) is introduced, leveraging decoupled horizontal and vertical strip convolutions within a grouped multi-branch architecture enhanced by an intuitionistic fuzzy channel attention mechanism. This design enables precise modeling of defect morphology and dominant orientation. Evaluated on benchmark datasets including NEU-Seg, the proposed method achieves a state-of-the-art mIoU of 89.60% with only 3.54 million parameters, outperforming current advanced approaches.
为解决OVCOS中视觉与文本语义鸿沟问题,提出ViCo-SAM3框架,通过视觉条件模块动态调整文本嵌入,增强跨模态交互。
为解决全景显著物体检测中的几何失真问题,提出TDFNet,通过三投影变形融合网络及交叉投影变形注意力模块提升检测性能。
该研究针对网络事件序列中的因果关系发现难题,提出了一种鲁棒的并发因果发现方法RCCD,通过引入注意力机制和自监督掩码重构框架来提高对缺失数据的鲁棒性和准确性。
This study addresses the prevalent issues of coarse, misaligned, or incomplete manual annotations in remote sensing semantic segmentation, which often distort model evaluation. To tackle this, the authors propose a training-free, reference-free mask fidelity assessment method that constructs counterfactual image pairs—preserving and erasing the region within the mask—and leverages a frozen vision-language model to evaluate whether class-specific evidence is concentrated inside the mask and absent outside it. This approach enables, for the first time, reference-free auditing of annotation quality in remote sensing segmentation, revealing systematic labeling biases across categories and facilitating automatic refinement of supervision signals. The proposed Contrastive Mask Fidelity (CMF) metric achieves 81% agreement with expert judgments across ten remote sensing datasets, substantially outperforming existing methods, and CMF-guided supervision significantly enhances cross-domain transfer performance.
This work addresses the challenge of effectively segmenting slender, anisotropic defects—such as cracks and scratches—on steel surfaces, which existing methods struggle to handle accurately. To this end, the authors propose the SPDCN network, which incorporates a Fuzzy-enhanced Multi-scale Context Module (FMCM) to adaptively fuse multi-scale contextual information. Additionally, an Adaptive Direction-Aware Deformable Convolution (ADADC) is introduced, leveraging decoupled horizontal and vertical strip convolutions within a grouped multi-branch architecture enhanced by an intuitionistic fuzzy channel attention mechanism. This design enables precise modeling of defect morphology and dominant orientation. Evaluated on benchmark datasets including NEU-Seg, the proposed method achieves a state-of-the-art mIoU of 89.60% with only 3.54 million parameters, outperforming current advanced approaches.