ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建ClearText-Video数据集,解决多模态大语言模型在处理真实世界文本中心视频时的质量敏感问题,涵盖视频恢复与场景文本增强任务。
📝 Abstract
Multimodal Large Language Models (MLLMs) have recently made strong progress in visual--linguistic understanding. However, their performance on text-centric video reasoning remains highly sensitive to input quality. Real-world user-provided videos often contain motion blur, compression artifacts, noise, and low-resolution text, which impair reliable text reading and downstream reasoning. Whether MLLMs can robustly read and reason about real-world scene text under diverse quality conditions remains a fundamental open question. We introduce ClearText-Video (CTVid), a large-scale, scene-text-aware benchmark for studying text-centric video understanding under controlled quality variation. CTVid contains 4,639 real-world text-rich egocentric videos, 550K+ frames, 1.6M human-verified scene-text annotations, and 220K+ spatial/temporal question--answer pairs in Chinese and English. For each high-quality video, CTVid provides content-matched Degraded-Quality and Restored-Quality variants, supporting two task families: Text-Centric Video Restoration and Multi-Quality VideoQA. We evaluate 18 representative restoration methods and 16 state-of-the-art MLLMs on CTVid. The results show that visual enhancement does not guarantee textual fidelity or downstream reasoning gains: blur is more damaging than low resolution, restored videos can alter the textual evidence used by MLLMs, and OCR-only pipelines remain far below direct multimodal reasoning. CTVid exposes the gap between video restoration and text-grounded understanding, providing a rigorous foundation for restoration-aware, quality-robust text-centric video systems.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Text-Centric Video Reasoning
Input Quality
Scene-Text Enhancement
Video Restoration
Innovation

Methods, ideas, or system contributions that make the work stand out.

ClearText-Video
Scene-Text Enhancement
Text-Centric Video Restoration
Multi-Quality VideoQA
Multimodal Large Language Models
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
J
Jinlong Li
OPPO US AI Center, USA
J
Jiaming Ding
OPPO US AI Center, USA
D
Dingfu Lu
OPPO US AI Center, USA; University of Wisconsin–Madison, USA
M
Malcolm Hsiu
OPPO US AI Center, USA; University of California San Diego, USA
C
Chuang Ke
OPPO US AI Center, USA
Kangning Yang
Kangning Yang
InnoPeak Technology
Deep LearningMutimodal LearningAffective ComputingHuman-Computer Interaction
B
Bochen Guan
OPPO US AI Center, USA
Lan Fu
Lan Fu
University of South Carolina
computer vision
Jie Cai
Jie Cai
OPPO AI Center
Computer VisionDeep Learning
Huiming Sun
Huiming Sun
OPPO US Research Center
Computer VisionRemote SensingSemantic Segmentation
Zibo Meng
Zibo Meng
OPPO
Computer VisionImage RestorationGenAI