Institution profile

WATRIX.AI

Industry researchasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

GaitSnippet: Gait Recognition Beyond Unordered Sets and Ordered Sequences

Aug 11, 2025

In gait recognition, unordered set modeling neglects short-term temporal dependencies, while ordered sequence modeling struggles to capture long-range correlations. To address this, we propose the “gait snippet” paradigm, representing human gait as a personalized, multi-scale composition of action snippets—thereby unifying short- and long-range temporal context modeling. Our method comprises two core components: snippet sampling and snippet modeling, leveraging a lightweight 2D convolutional backbone for efficient snippet-level feature extraction and aggregation. This work introduces the snippet concept to gait recognition for the first time, breaking away from the conventional dichotomy of set- versus sequence-based modeling. Evaluated on Gait3D and GREW benchmarks, our approach achieves rank-1 accuracies of 77.5% and 81.7%, respectively, demonstrating strong effectiveness, robustness, and cross-scenario generalization capability.

0 citationsRead paper

OmniDiff: A Comprehensive Benchmark for Fine-grained Image Difference Captioning

Mar 14, 2025

Existing Image Difference Captioning (IDC) datasets suffer from limited scope—narrow scene coverage—and shallow granularity—coarse-grained descriptions—hindering fine-grained understanding in complex, dynamic environments. To address this, we introduce OmniDiff, the first fine-grained IDC benchmark spanning both real-world and 3D-synthetic scenes, comprising 324 diverse scenarios, 12 change categories, and human-annotated captions averaging 60 words each. Methodologically, we propose a plug-and-play Multi-scale Difference Perception (MDP) module and build M$^3$Diff, an end-to-end multimodal large model integrating visual encoding, cross-modal alignment, and MDP. Our approach achieves state-of-the-art performance across five benchmarks—including Spot-the-Diff and CLEVR-Change—with significant improvements in cross-scenario difference recognition accuracy. All data, code, and models are publicly released.

0 citationsRead paper
Recent publications

Latest Papers

GaitSnippet: Gait Recognition Beyond Unordered Sets and Ordered Sequences

Aug 11, 2025

In gait recognition, unordered set modeling neglects short-term temporal dependencies, while ordered sequence modeling struggles to capture long-range correlations. To address this, we propose the “gait snippet” paradigm, representing human gait as a personalized, multi-scale composition of action snippets—thereby unifying short- and long-range temporal context modeling. Our method comprises two core components: snippet sampling and snippet modeling, leveraging a lightweight 2D convolutional backbone for efficient snippet-level feature extraction and aggregation. This work introduces the snippet concept to gait recognition for the first time, breaking away from the conventional dichotomy of set- versus sequence-based modeling. Evaluated on Gait3D and GREW benchmarks, our approach achieves rank-1 accuracies of 77.5% and 81.7%, respectively, demonstrating strong effectiveness, robustness, and cross-scenario generalization capability.

0 citationsRead paper

OmniDiff: A Comprehensive Benchmark for Fine-grained Image Difference Captioning

Mar 14, 2025

Existing Image Difference Captioning (IDC) datasets suffer from limited scope—narrow scene coverage—and shallow granularity—coarse-grained descriptions—hindering fine-grained understanding in complex, dynamic environments. To address this, we introduce OmniDiff, the first fine-grained IDC benchmark spanning both real-world and 3D-synthetic scenes, comprising 324 diverse scenarios, 12 change categories, and human-annotated captions averaging 60 words each. Methodologically, we propose a plug-and-play Multi-scale Difference Perception (MDP) module and build M$^3$Diff, an end-to-end multimodal large model integrating visual encoding, cross-modal alignment, and MDP. Our approach achieves state-of-the-art performance across five benchmarks—including Spot-the-Diff and CLEVR-Change—with significant improvements in cross-scenario difference recognition accuracy. All data, code, and models are publicly released.

0 citationsRead paper