SignDino: Self-Supervised Sign Language Representation Learning via Temporal-Axis Self-Distillation

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出SignDino,通过将DINOv3的自监督学习方法从图像空间域转移到手语视频的时间域,解决了手语表征学习中特有的时序和解剖结构问题。
📝 Abstract
Self-supervised sign language representation learning must model two properties not central to natural-image SSL: signs are produced by a small set of anatomically distinct articulators, and their meaning depends on the temporal organisation of those articulators. We introduce SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams. Each video is decomposed into left-hand, right-hand, and face streams by a detector-first YOLOv8n+ByteTrack pipeline. A frozen DINOv3 ViT-B/16 embeds each per-frame anatomical crop, while lightweight temporal Transformers, not the image backbone, form the student and EMA teacher. They are trained by temporal DINO self-distillation, frame-level masked-token prediction in the style of iBOT, KoLeo feature spreading, and Gram anchoring of the frame-to-frame similarity structure. This design keeps strong image-level visual primitives fixed and learns only how articulator states evolve across time. We evaluate on sign-to-English translation, isolated sign recognition, and fingerspelling detection benchmarks. Across these tasks, SignDino provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evaluation.
Problem

Research questions and friction points this paper is trying to address.

Self-Supervised
Sign Language
Temporal Organization
Articulators
Representation Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Temporal-Axis Self-Distillation
Self-Supervised Learning
Sign Language Representation
Temporal Transformers
Articulator State Evolution
J
Junyi Hu
New York University Abu Dhabi
Z
Zhewen He
New York University Abu Dhabi
H
Haomian Huang
New York University Abu Dhabi
Z
Zhenhua Li
New York University Abu Dhabi
Zhifei Li
Zhifei Li
Research Scientist at Google
machine translationnatural language processingmachine learningwireless networks
Y
Yi Fang
New York University Abu Dhabi