SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决连续手语识别中的细粒度表示学习和视频-文本对齐效率问题,提出SMART框架,利用MLLM生成的运动描述作为辅助语义线索,并引入多尺度时间适配器和CSFormer模块。
📝 Abstract
Continuous sign language recognition (CSLR) aims to recognize gloss sequences from unsegmented sign videos under weak sequence-level supervision. However, existing methods rely on sentence-level gloss annotations, providing limited temporal and semantic guidance for fine-grained representation learning. Conventional video-text alignment also requires large batch sizes, making it inefficient for memory-intensive sign language video training. In this work, we propose SMART, an MLLM-guided temporal alignment framework for joint sign recognition and spotting. SMART uses MLLMgenerated motion descriptions as auxiliary semantic cues and performs stable videotext alignment under small-batch training. To improve temporal representation learning, we introduce a Multi-Scale Temporal Adapter that models temporal interactions during transformer encoding. For dense temporal localization, SMART incorporates CSFormer, a CSLR-guided spotting module that injects recognition-derived gloss evidence into a boundary-aware spotting network. This unified framework enables CSLR features to benefit spotting, while spotting supervision complements weak CTC-based recognition. Experiments on four sign language benchmarks, including PHOENIX14-T, CSL-Daily, Large-scale KSL, and Disaster and Safety KSL datasets, demonstrate the effectiveness of SMART across both recognition and spotting tasks.
Problem

Research questions and friction points this paper is trying to address.

continuous sign language recognition
temporal alignment
fine-grained representation learning
video-text alignment
small-batch training
Innovation

Methods, ideas, or system contributions that make the work stand out.

MLLM-guided temporal alignment
Multi-Scale Temporal Adapter
CSFormer
small-batch training
unified framework for recognition and spotting
E
Eunjee Choi
Dankook University, Yongin, Korea
J
JungHoon Sung
Dankook University, Yongin, Korea
S
Seongwhan Cho
Dankook University, Yongin, Korea
C
Chu Xin
Dankook University, Yongin, Korea
Y
Younggeun Choi
Dankook University, Yongin, Korea