MLLM-Guided Semantic Correction for Text-to-Video Generation

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses semantic deviation in text-to-video generation by proposing a training-free mid-course correction framework. The approach pioneers the integration of multimodal large language model feedback into the diffusion sampling loop, employing a semantic evaluation supervisor to diagnose intermediate frame discrepancies and a modification assistant for controllable latent space intervention. This mechanism enables dynamic self-reflection and trajectory rectification during inference. Experimental results demonstrate that, without parameter fine-tuning, the proposed method significantly enhances semantic alignment, visual fidelity, and temporal consistency across multiple benchmarks. Consequently, this work establishes a novel, interpretable paradigm for semantic correction within the generative process, offering a robust solution to persistent alignment challenges in video synthesis.
📝 Abstract
Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, these models often suffer from semantic errors such as missing objects, incorrect attributes, or mismatched actions. Although some semantic correction methods perform optimization before sampling or refinement after sampling, how to detect and correct semantic deviations during the video generation process remains underexplored. In this paper, we introduce a training-free, interpretable mid-generation correction framework that integrates multimodal large language model (MLLM) feedback directly into the diffusion sampling loop. Our framework achieves diffusion trajectory correction by injecting semantic evaluation signals during video synthesis, enabling the model to optimize the generated content through continuous self-reflection. We propose two key modules: a Semantic Assessment Supervisor that generates intermediate preview frames for semantic evaluations and deviation diagnostics, and a Semantic Modification Assistant that corrects semantic drift during inference via a controllable latent trajectory intervention. Our method improves semantic alignment, visual fidelity, and temporal consistency without modifying model parameters. We validate the effectiveness of our approach through extensive experiments across multiple benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Text-to-Video Generation
Semantic Errors
Semantic Correction
Mid-generation Correction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-free Mid-generation Correction
MLLM-Guided Feedback
Diffusion Trajectory Intervention
Semantic Assessment Supervisor
Self-reflection
J
Junhao Chen
School of Software Technology, Zhejiang University, Hangzhou 310027, China
Z
Zheqi Lv
College of Computer Science and Technology, Zhejiang University, Hangzhou 310027, China
K
Keting Yin
School of Software Technology, Zhejiang University, Hangzhou 310027, China
S
Shengyu Zhang
College of Computer Science and Technology, Zhejiang University, Hangzhou 310027, China
Zhou Zhao
Zhou Zhao
Zhejiang University
Machine LearningData MiningMultimedia Computing
F
Feiyang Chen
AI System Innovation Lab, Huawei Cloud Computing Technology Co., Ltd., Hangzhou 310005, China
Xinyu Duan
Xinyu Duan
Huawei Cloud
LLMInference Optimization
Baoxing Huai
Baoxing Huai
HuaweiCloud
NLPKnowledge Computing
F
Fei Wu
College of Computer Science and Technology, Zhejiang University, Hangzhou 310027, China