Who Remains, What Changes: Identity Anchored Composed Gait Retrieval

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种基于自然语言指令和参考序列的目标步态检索方法ComposeGait,通过身份锚定的组合框架解决身份漂移问题。
📝 Abstract
Gait recognition has achieved remarkable progress, yet existing methods remain confined to rigid visual matching and often overlook the potential of natural language instructions for interactive retrieval. In this paper, we introduce Composed Gait Retrieval (CoGR), a novel task that retrieves a target gait sequence based on a reference sequence and a natural language modification query. To address the absence of existing datasets for this task, we design an automated annotation pipeline powered by large vision-language models (VLMs) to construct the first gait-language datasets: Language-Augmented CCPG and Language-Augmented CASIA-B. Building on this, we propose ComposeGait, an identity-anchored composition framework designed to prevent the identity drift that arises when generic composed retrieval follows the instruction but returns the wrong person. Its Part-aware Identity Adapter (PIA) aggregates multi-frame, part-aware identity evidence into a sample-specific ID token. We inject the ID tokens into both branches of a shared Q-Former to preserve identity, while excluding the ID-token outputs from the final retrieval embeddings. Joint identity and task-adapted composed-retrieval objectives optimize this space end to end. We evaluate ComposeGait on both benchmarks and show that it achieves the best R@1 among the compared methods, reaching 72.38% on Language-Augmented CCPG and 83.61% on Language-Augmented CASIA-B. These results establish ComposeGait as a strong baseline for CoGR. The datasets and code will be made publicly available.
Problem

Research questions and friction points this paper is trying to address.

Gait Recognition
Natural Language Instructions
Interactive Retrieval
Composed Gait Retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Composed Gait Retrieval
Identity-Anchored Composition
Part-aware Identity Adapter
Natural Language Instructions
Automated Annotation Pipeline
J
Jingchen Fei
Beijing University of Posts and Telecommunications
Z
Zengbin Wang
Beijing University of Posts and Telecommunications
Y
Yukun Liu
Huazhong University of Science and Technology
Muyi Sun
Muyi Sun
School of AI, BUPT (<< NLPR CASIA << BUPT)
Multi-Modality LearningComputer VisionBiometricsMedical Image Analysis
Shibiao Xu
Shibiao Xu
Beijing University of Posts and Telecommunications
Computer VisionMachine LearningComputer Graphics
M
Man Zhang
Beijing University of Posts and Telecommunications