MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval
Existing action–text retrieval benchmarks are limited by action homogeneity, imbalanced category distributions, and oversimplified textual descriptions, hindering the evaluation of cross-domain and cross-granularity alignment capabilities. To address these limitations, this work proposes MRBench—a comprehensive benchmark featuring heterogeneous actions, balanced category distribution, and multi-granularity textual descriptions—alongside a lightweight granularity-aware model. Leveraging multi-source action data and large language models, the authors construct 3,390 action instances paired with 10,170 multi-granularity text descriptions. The model employs a frozen backbone with branch-specific adapters and a comparability-preserving score fusion strategy. Experiments reveal that existing methods exhibit significant generalization gaps and granularity sensitivity on MRBench, whereas the proposed approach substantially improves fine-grained retrieval performance while maintaining competitive standard retrieval accuracy.