SignRefine: Adapting Foundational Video Models for Sign Language Generation

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决现有视频生成模型在手语视频生成中的不准确问题,提出SignRefine模型,通过局部适配器改进手部和面部区域的精确度,并使用大规模自然手语视频数据集NVSign进行训练。
📝 Abstract
Sign language video generation demands precise hand and facial articulation, yet modern video diffusion models, trained predominantly on spoken-language video, produce artifacts that render signing unintelligible. We propose SignRefine, a sign language video generation model that produces comprehensible signing from 2D keypoint conditioning alone, generalizing across appearances and visual conditions. Our approach builds on a pretrained video diffusion transformer and introduces local adapters with spatial grounding to selectively refine hand and face regions, steering the strong base model's prior toward accurate articulation. To enable this work and support broader sign language research, we present NVSign, a large-scale dataset of video content natively produced in sign language, offering diverse signer appearances, environments, and natural conversational settings. Trained on this data, our model shows up to 30% improvement in hand pose precision metrics over the strongest baseline and is preferred by sign language users for visual quality and comprehensibility in more than 80% of comparisons.
Problem

Research questions and friction points this paper is trying to address.

Sign Language Generation
Video Diffusion Models
Articulation Precision
Innovation

Methods, ideas, or system contributions that make the work stand out.

SignRefine
local adapters with spatial grounding
NVSign dataset
hand and facial articulation refinement