From Visual to Multimodal: Systematic Ablation of Encoders and Fusion Strategies in Animal Identification
This study addresses the limitations of existing pet identification systems, which suffer from small-scale datasets and reliance on a single visual modality, leading to suboptimal performance in lost-pet retrieval tasks. To overcome these challenges, the authors construct a large-scale pet dataset comprising 1.9 million images and present the first systematic exploration of multimodal fusion methods in this domain. They propose an enhanced framework that incorporates synthetic text descriptions as semantic priors, leveraging the SigLIP2-Giant visual encoder and the E5-Small-v2 text encoder. Through comprehensive evaluation of fusion strategies—ranging from feature concatenation to adaptive gating—the study demonstrates that gated fusion significantly sharpens decision boundaries. Experimental results show a Top-1 accuracy of 84.28% and an equal error rate of 0.0422, representing an 11% improvement over the best single-modality baseline.