Environment-Invariant Subspace Learning for Generalizable Deepfake Detection
为解决深伪检测中的环境干扰问题,提出环境不变子空间学习框架EISL,通过可学习低秩投影分解特征,增强模型对未知伪造类型和环境变化的鲁棒性。
为解决深伪检测中的环境干扰问题,提出环境不变子空间学习框架EISL,通过可学习低秩投影分解特征,增强模型对未知伪造类型和环境变化的鲁棒性。
为解决热带气旋预报不确定性问题,提出基于物理约束的生成式AI框架Tianmu-TC,有效提高预报准确性和效率。
This work addresses the challenge of camouflaged object detection, where high similarity between foreground and background and weak boundary cues hinder accurate segmentation. To this end, the authors propose the LAD-COD framework, which introduces a novel Language-Aligned Dual Visual Fusion (LADVF) mechanism. This mechanism propagates language-instruction-guided semantic information to the image patch level and aligns it with low-level dense visual features, enabling synergistic optimization between semantic guidance and fine structural perception. Additionally, a trainable hierarchical visual branch is incorporated to extract camouflage-sensitive features, and a residual gating mechanism fuses multi-source information effectively. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance across all 12 dataset–metric combinations on three major benchmarks: CAMO, COD10K, and NC4K.
This work addresses the limitation of existing vision foundation model–based remote sensing change detection methods, which process bi-temporal images independently and thus fail to effectively model cross-temporal relationships. To overcome this, the authors propose AdaDINO, a framework that introduces a bi-temporal interaction mechanism into a frozen DINO backbone by coupling dual streams at selected layers and injecting shared temporal residuals with opposite signs to explicitly capture temporal dependencies. Key innovations include Change-aware Gated Local Adaptation (CGLA), Batch-shared Block Selection (BSCS), and a CGLA Prior-guided Refinement (CPGR) decoder, enabling efficient computation while keeping the backbone frozen. Experiments demonstrate state-of-the-art performance across four benchmarks, achieving an F1 score of 85.29% on SYSU-CD, a 62.5% reduction in FFN width, and a 1.41× throughput improvement.
This work addresses the challenges of performance degradation and high computational cost in multimodal crack segmentation for industrial facilities under arbitrary modality missingness. The authors propose Compass, a lightweight network featuring Degradation Simulation Distillation (DSD), a Modality-Agnostic Feature-Aware Prototype Transformer (FAPT), Evidential Theory-driven Topology-Preserving Fusion (ETPF), and an uncertainty-gated decoder to achieve robust segmentation regardless of missing modalities. Built upon an efficient Needle RWKV backbone, Compass contains only 2.58 million parameters. It achieves state-of-the-art performance across three datasets, with CrackDepth attaining an F1 score of 0.8216 and mIoU of 0.8434 even when 90% of depth modality data is missing.
为解决深伪检测中的环境干扰问题,提出环境不变子空间学习框架EISL,通过可学习低秩投影分解特征,增强模型对未知伪造类型和环境变化的鲁棒性。
为解决热带气旋预报不确定性问题,提出基于物理约束的生成式AI框架Tianmu-TC,有效提高预报准确性和效率。
This work addresses the challenge of camouflaged object detection, where high similarity between foreground and background and weak boundary cues hinder accurate segmentation. To this end, the authors propose the LAD-COD framework, which introduces a novel Language-Aligned Dual Visual Fusion (LADVF) mechanism. This mechanism propagates language-instruction-guided semantic information to the image patch level and aligns it with low-level dense visual features, enabling synergistic optimization between semantic guidance and fine structural perception. Additionally, a trainable hierarchical visual branch is incorporated to extract camouflage-sensitive features, and a residual gating mechanism fuses multi-source information effectively. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance across all 12 dataset–metric combinations on three major benchmarks: CAMO, COD10K, and NC4K.
This work addresses the limitation of existing vision foundation model–based remote sensing change detection methods, which process bi-temporal images independently and thus fail to effectively model cross-temporal relationships. To overcome this, the authors propose AdaDINO, a framework that introduces a bi-temporal interaction mechanism into a frozen DINO backbone by coupling dual streams at selected layers and injecting shared temporal residuals with opposite signs to explicitly capture temporal dependencies. Key innovations include Change-aware Gated Local Adaptation (CGLA), Batch-shared Block Selection (BSCS), and a CGLA Prior-guided Refinement (CPGR) decoder, enabling efficient computation while keeping the backbone frozen. Experiments demonstrate state-of-the-art performance across four benchmarks, achieving an F1 score of 85.29% on SYSU-CD, a 62.5% reduction in FFN width, and a 1.41× throughput improvement.
This work addresses the challenges of performance degradation and high computational cost in multimodal crack segmentation for industrial facilities under arbitrary modality missingness. The authors propose Compass, a lightweight network featuring Degradation Simulation Distillation (DSD), a Modality-Agnostic Feature-Aware Prototype Transformer (FAPT), Evidential Theory-driven Topology-Preserving Fusion (ETPF), and an uncertainty-gated decoder to achieve robust segmentation regardless of missing modalities. Built upon an efficient Needle RWKV backbone, Compass contains only 2.58 million parameters. It achieves state-of-the-art performance across three datasets, with CrackDepth attaining an F1 score of 0.8216 and mIoU of 0.8434 even when 90% of depth modality data is missing.