MMGait: Benchmarking and Unifying Gait Recognition across Heterogeneous Modalities

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入MMGait多传感器基准和Omni-Modal Gait Recognition方法,解决了跨异构模态步态识别的问题。
📝 Abstract
Gait recognition is commonly studied using RGB videos or their derived silhouettes and poses. Yet human walking produces heterogeneous photometric, geometric, and motion cues that cannot be systematically examined with RGB-centered benchmarks. We present MMGait, a large-scale multi-sensor benchmark that brings visible, infrared, depth, LiDAR, and radar observations into sequence-level correspondence. It provides diverse modalities spanning appearance, contours, geometry, motion, and body structure. Under a shared impostor-augmented protocol, we evaluate single-modal recognition, cross-modal recognition via directed retrieval, and multi-modal recognition using task-specific experts. Across settings, modality rankings vary with probe conditions, cross-modal alignment remains difficult, and fusion often provides complementary gains. This analysis exposes a scalability problem: individual modalities, modality pairs, and fusion configurations are typically handled by separately trained experts. We formulate Omni-Modal Gait Recognition, which unifies single-modal, cross-modal, and multi-modal recognition within a shared identity space. OmniGait++ uses modality-specific front ends followed by a shared identity encoder to preserve modality-dependent cues while learning comparable identity descriptors. An anchor-guided fusion module aggregates modality subsets of varying size without frame-level synchronization. A jointly trained checkpoint covers all three recognition settings and accommodates modality subsets of different compositions and cardinalities. Experiments show OmniGait++ remains competitive with task-specific experts in many shared settings and extends to higher-cardinality fusion unavailable to fixed-pair models. The results establish MMGait as a common testbed for heterogeneous gait sensing and demonstrate the feasibility of unified recognition under varying modality availability.
Problem

Research questions and friction points this paper is trying to address.

Gait Recognition
Heterogeneous Modalities
Scalability Problem
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-sensor benchmark
Omni-Modal Gait Recognition
Modality-specific front ends
Shared identity encoder
Anchor-guided fusion
💼 Related Jobs
No related jobs found.
Saihui Hou
Saihui Hou
Beijing Normal University
Deep LearningComputer VisionMultimodal Large Language Models
Chenye Wang
Chenye Wang
M.E., Beijing Normal University
Q
Qingyuan Cai
School of Artificial Intelligence, Beijing Normal University, Beijing, China
A
Aoqi Li
JD.com, Beijing, China (also formerly with School of Artificial Intelligence, Beijing Normal University)
Yongzhen Huang
Yongzhen Huang
School of Artificial Intelligence, Beijing Normal University
Computer VisionPattern RecognitionDeep Learning