3D-USE: From Image-Level to Scene-Level Underwater Enhancement
本文提出3D-USE框架,通过两阶段方法从降质的多视角水下图像中学习可见度增强的3D场景表示,解决水下3D重建中的颜色偏移和可见度损失问题。
本文提出3D-USE框架,通过两阶段方法从降质的多视角水下图像中学习可见度增强的3D场景表示,解决水下3D重建中的颜色偏移和可见度损失问题。
Existing vision foundation models lack a backward guidance mechanism whereby high-level reasoning informs low-level perceptual feature learning, resulting in an incomplete perception–reasoning–generation loop. Method: We propose Vision Instruction Pre-training (ViTP), a novel paradigm built upon Vision Transformers and trained end-to-end on large-scale, domain-specific vision–language instruction data via instruction following and cross-modal alignment. Crucially, we introduce Vision Robustness Learning (VRL), which explicitly encourages the model to extract domain-relevant, interference-resistant, discriminative features from sparse visual tokens. Contribution/Results: ViTP establishes the first end-to-end, domain-specialized foundation model capable of reasoning-driven perceptual optimization. Evaluated on 16 remote sensing and medical imaging benchmarks, it achieves state-of-the-art performance, significantly improving generalization and semantic understanding across downstream tasks.
本文提出3D-USE框架,通过两阶段方法从降质的多视角水下图像中学习可见度增强的3D场景表示,解决水下3D重建中的颜色偏移和可见度损失问题。
Existing vision foundation models lack a backward guidance mechanism whereby high-level reasoning informs low-level perceptual feature learning, resulting in an incomplete perception–reasoning–generation loop. Method: We propose Vision Instruction Pre-training (ViTP), a novel paradigm built upon Vision Transformers and trained end-to-end on large-scale, domain-specific vision–language instruction data via instruction following and cross-modal alignment. Crucially, we introduce Vision Robustness Learning (VRL), which explicitly encourages the model to extract domain-relevant, interference-resistant, discriminative features from sparse visual tokens. Contribution/Results: ViTP establishes the first end-to-end, domain-specialized foundation model capable of reasoning-driven perceptual optimization. Evaluated on 16 remote sensing and medical imaging benchmarks, it achieves state-of-the-art performance, significantly improving generalization and semantic understanding across downstream tasks.