SignMatch: Matching Dictionary Signs to Continuous Sign Language Video

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文旨在通过学习连续手语视频的标志嵌入空间来匹配字典标志与连续手语视频中的相应标志,利用视觉相似性实现直接匹配,并展示出跨数据集和手语类型的强泛化能力。
📝 Abstract
The objective of this paper is to match dictionary sign videos to corresponding signs in continuous signing videos, where a match is defined by the visual similarity alone - the handshape and motion relative to the body. To achieve this, we learn a prototype-structured sign embedding space from continuous video annotated with signs, where each learnable prototype corresponds to a sign class. Isolated dictionary videos are then mapped into this sign space, enabling the matching between dictionary exemplars and continuous sign instances. This design supports direct dictionary-guided sign matching through embedding similarity and naturally extends to unseen signs using only dictionary exemplars. Experiments on ASL-Citizen dictionary retrieval, ChaLearn OSLWL dictionary-to-continuous sign matching, and using BOBSL's CSLR2 evaluation for automatic sign annotation demonstrate strong generalisation across datasets, tasks and sign languages. Without benchmark-specific supervision, the learned representation transfers effectively across American, British, and Spanish Sign Languages, outperforming prior methods on all three benchmarks. Project page: https://www.robots.ox.ac.uk/~vgg/research/signmatch/
Problem

Research questions and friction points this paper is trying to address.

sign matching
continuous sign language video
dictionary signs
visual similarity
Innovation

Methods, ideas, or system contributions that make the work stand out.

prototype-structured sign embedding space
dictionary-guided sign matching
generalisation across datasets and sign languages
💼 Related Jobs
No related jobs found.
R
Ryan Wong
Visual Geometry Group, Department of Engineering Science, University of Oxford, UK
Youngjoon Jang
Youngjoon Jang
KAIST
Computer VisionMachine Learning
Liliane Momeni
Liliane Momeni
University of Oxford
Computer VisionMachine LearningArtificial Intelligence
G
Gül Varol
Visual Geometry Group, Department of Engineering Science, University of Oxford, UK; LIGM, École des Ponts, IP Paris, Univ Gustave Eiffel, CNRS, France
Andrew Zisserman
Andrew Zisserman
University of Oxford
Computer VisionMachine Learning