Wildlife Target Re-Identification Using Self-supervised Learning in Non-Urban Settings
To address the scarcity of labeled data for wildlife re-identification in non-urban environments, this paper proposes a temporal self-supervised learning framework tailored for camera-trap videos. The method leverages unlabeled consecutive video frames to model temporal consistency of individual appearance and employs contrastive learning to extract robust, view-invariant feature representations. Its key contribution is the first systematic integration of temporal self-supervision into open-world wildlife re-identification—enabling discriminative individual representation learning without manual annotations. Experiments demonstrate that the approach significantly outperforms supervised baselines across multiple species (e.g., leopard cats, wild boars), particularly excelling in few-shot and cross-domain generalization. Moreover, the learned features achieve state-of-the-art performance on downstream tasks including image retrieval, unsupervised clustering, and few-shot fine-tuning.