SapiensID 2.0: Aligning Human Recognition Foundation Models with Human Perception

๐Ÿ“… 2026-08-11
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing human recognition models rely heavily on static geometric features, overlooking the semantic and temporal mechanisms inherent in human perception. This limitation leads to semantic blind spots, heightened sensitivity to transient noise, and an inability to effectively leverage soft biometric cues and motion dynamics. To address these issues, this work proposes a perception-aligned recognition framework that integrates multimodal large language models for zero-shot semantic transfer. The framework introduces two key components: Invariant Trait Alignment (ITA) and Transient Noise Disentanglement (TND), along with a novel Kinematic Semantic Attention Head (K-SAH) that captures motion semantics across temporal windowsโ€”without requiring large-scale video data. The approach achieves state-of-the-art performance on person re-identification and gait recognition tasks while maintaining robust face recognition capabilities.
๐Ÿ“ Abstract
While foundation models have significantly advanced human recognition across diverse modalities, they predominantly rely on static, geometric feature extraction. This approach fundamentally diverges from human perception. Consequently, current models often suffer from "semantic blindness," overfitting to transient noise while failing to leverage invariant soft biometrics, and struggle to capture temporal motion signatures. To bridge this gap, we propose SapiensID 2.0, a human recognition framework enriched with both semantic and temporal awareness. To overcome the lack of soft-biometric annotations, we transfer zero-shot semantic knowledge from Multimodal Large Language Models (MLLMs) into a discriminative embedding space. We resolve the dimensional mismatch between these spaces using Invariant Trait Alignment (ITA) to distill core persistent traits, and Transient Noise Disentanglement (TND) to decouple artifacts like clothing. Furthermore, we design a Kinematic Semantic Attention Head (K-SAH) that extends spatial attention across temporal windows. By tracking semantic patches over time, K-SAH captures rich kinematic signatures without requiring large-scale video datasets. Extensive experiments demonstrate that SapiensID 2.0 achieves state-of-the-art performance across image- and video-based person re-identification and gait recognition, while maintaining robust face recognition capabilities.
Problem

Research questions and friction points this paper is trying to address.

human recognition
semantic blindness
soft biometrics
temporal motion signatures
foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Invariant Trait Alignment
Transient Noise Disentanglement
Kinematic Semantic Attention Head
Zero-shot Semantic Transfer
Temporal Motion Signatures