HVPNet: A Bio-Inspired Network for General Salient and Camouflaged Object Detection
Existing multimodal salient and camouflaged object detection methods suffer from structural complexity and large parameter counts, making it challenging to balance accuracy and efficiency. Inspired by the human visual system, this work proposes a lightweight, unified architecture that integrates a Retinal Integration Module (RIM) for hierarchical, multi-stage cross-modal feature fusion and a Cortical Decoder (CD) that mimics visual cortical mechanisms for layered decoding. This approach establishes a biologically inspired, simplified modeling paradigm capable of supporting diverse modalities and tasks within a single framework. Evaluated across four modalities, seven tasks, and 22 datasets, the model achieves an excellent trade-off between accuracy and efficiency with a compact structure, demonstrating strong generalization capability.