RAIDAL: Redundancy-Aware Information Density Active Learning for CTC-Based Continuous Sign Language Recognition

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对连续手语识别中注释成本高的问题,提出RAIDAL方法,利用CTC解码器识别关键帧,提高主动学习效率。
📝 Abstract
Continuous sign language recognition (CSLR) is a key technology for accessibility, yet its development remains limited by the high cost of annotating continuous video streams. Active learning offers a path toward mitigating this cost, but standard acquisition functions are not designed for weakly aligned sign language videos, where sign executions are interleaved with rest poses, irregular pauses, sign-like motion, and temporally redundant frames. This temporal redundancy can undermine sample selection, as acquisition scores may be influenced by timesteps from regions that are not associated with the decoded gloss sequence, distorting the video's estimated informativeness. In this work, we show that modern CSLR models already contain a mechanism for identifying gloss-level temporal evidence: the CTC decoder. Although typically used only during inference, its alignment peaks indicate where the model localizes each predicted gloss in the feature sequence, providing a source of temporal structure for active learning acquisition functions at zero additional labeling cost. Thus, we introduce RAIDAL (Redundancy-Aware Information Density Active Learning), which repurposes the CTC decoder to restrict representation-based scoring to decoder-aligned gloss regions, rather than exposing the acquisition function to the entire unfiltered video. Across three datasets and two architectures, RAIDAL achieves its strongest data-efficiency gains over competing baselines in large-vocabulary, budget-limited settings, while remaining competitive in the smaller-vocabulary, large-budget setting. The code used in this work is publicly available at github.com/verlab/RAIDAL.
Problem

Research questions and friction points this paper is trying to address.

Continuous Sign Language Recognition
Active Learning
Temporal Redundancy
Annotation Cost
CTC Decoder
Innovation

Methods, ideas, or system contributions that make the work stand out.

Redundancy-Aware
Information Density
Active Learning
CTC Decoder
Continuous Sign Language Recognition
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Rafael A. Diniz Augusto
Universidade Federal de Minas Gerais
G
Gabriel L. Oliveira
University of Bristol
Erickson R. Nascimento
Erickson R. Nascimento
Universidade Federal de Minas Gerais