SLAP: Selective Local Vision-Language Alignment for Fish Re-Identification via Partial Optimal Transport
This work addresses the limitations of existing CLIP-based methods in fish re-identification, which rely on global image-text alignment and are thus susceptible to background clutter and non-discriminative regions. To overcome this, the authors propose a selective local visual-language alignment framework that, for the first time, incorporates partial optimal transport (POT) into fine-grained fish re-identification. By leveraging POT, the method establishes selective correspondences between local image patch embeddings and multiple identity-aware textual prompts, thereby enhancing strongly relevant cross-modal associations while suppressing interference from irrelevant regions. Built upon the CLIP architecture and integrating local embedding, multi-prompt encoding, and end-to-end training, the proposed approach significantly outperforms existing CLIP-based methods on the Symphodus melops dataset and demonstrates superior generalization across multiple marine species re-identification benchmarks.