Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入Arti-JEPA模型,解决了实时MRI声带图像分析中标注数据稀缺和模态差异问题,采用自监督学习方法,并在多个任务上验证了其有效性。
📝 Abstract
Real-time MRI (rtMRI) captures the dynamics of the entire vocal tract during speech, but labeled data are scarce and the modality - single-slice, grayscale, low-resolution - differs substantially from the natural videos that video foundation models are trained on. We introduce Arti-JEPA, a joint embedding predictive architecture to model vocal tract rtMRI by continuing its self-supervised objective on about 62h of unlabelled vocal-tract videos, and evaluate the frozen representation on three tasks: cross-domain phoneme prediction (on typical speakers), fluent-vs-disfluent classification (a corpus containing stuttered speech), and characterizing pre/post-operative transfer (after partial glossectomy). Three key findings emerge. (1) A temporal video prior decisively outperforms per-frame image encoders, and latent prediction (V-JEPA) is at least as strong as pixel reconstruction (VideoMAE), with the edge on fine-grained phonemes. (2) Domain adaptation is \emph{task-dependent}: it roughly doubles cross-domain phoneme prediction $κ$ (to 0.352) but does not help binary stuttering classification. (3) Arti-JEPA was able to recover phoneme signal from pre/post glossectomy speech --- an in-domain probe decodes patients at least as well as a typical speaker, indicating that the residual transfer gap is cross-speaker/domain misalignment, not surgical signal loss, and post-operative decoding does not fall below performance on pre-operative speech. Together, these position a frozen, domain-adapted rtMRI encoder as a reusable measurement tool for articulatory and clinical speech science.
Problem

Research questions and friction points this paper is trying to address.

real-time MRI
vocal tract
unlabelled data
speech-production analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Arti-JEPA
Joint Embedding Predictive Architecture
Real-time MRI (rtMRI)
Temporal Video Prior
Domain Adaptation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Hong Nguyen
Hong Nguyen
PhD Student at University of Southern California
Video UnderstandingMultimodality ModelsHuman-centric AIBehavioural Models
Sean Foley
Sean Foley
Macquarie University
Applied FinanceDigital FinanceMarket MicrostructureCryptocurrenciesDeFi
C
Christina Hagedorn
Department of English, College of Staten Island, City University of New York, New York, NY, USA
Y
Yijing Lu
Department of Linguistics, University of Potsdam, Germany
Sudarsana Reddy Kadiri
Sudarsana Reddy Kadiri
University of Southern California
Speech ProcessingBiomedical SignalsMultimodalityHealthcare InformaticsDeep Learning
D
Dani Byrd
Department of Linguistics, University of Southern California, Los Angeles, CA, USA
S
Shrikanth Narayanan
Ming Hsieh Department of Electrical and Computer Engineering, University of Southern California, Los Angeles, CA, USA