MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究揭示视频大语言模型无法准确理解运动,并通过MotionBlind基准测试验证,该测试使用成对的自录视频来评估模型在速度、大小和方向上的物理基础运动理解能力。
📝 Abstract
Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: how fast something moves, which way it travels, how hard it is pushed. We show they cannot. A Video-LLM can watch two clips of the same person in the same room, name every object in both, and still fail to say which clip moves faster. We introduce MotionBlind, a contrastive benchmark of self-recorded video for physically grounded motion(speed, magnitude, and direction), the variables a world model must predict. Each instance is a pair of near-identical clips that differ only in motion. Each clip carries two complementary yes/no questions, giving four items per instance, and a model earns credit only if all four are correct. We report Instance Accuracy(IAcc), which has a 6.25% chance floor. Single-frame, appearance, and language-only shortcuts all collapse to it. MotionBlind complements the recent TimeBlind benchmark. We run a controlled study of six open and two frontier Video-LLMs, varying whether the video is present, whether frames are shown in the correct temporal order, and how frames are sampled (1 to 24 frames, four selection strategies). Open models sit near the 6.25% floor, and scale does not help. Removing the video drops every model to zero IAcc, and shuffling frames collapses IAcc to chance, so the task genuinely needs video in order. Neither more frames nor smarter frame selection closes the gap, because these change which frames are seen, not whether motion is read. Only Gemini3.1 Pro clears the benchmark overall, and even it fails on speed. A frontend that cannot tell two speeds of the same action apart is not yet a trustworthy source of supervision, reward, or evaluation for a world model.
Problem

Research questions and friction points this paper is trying to address.

Video-LLMs
motion understanding
world models
Innovation

Methods, ideas, or system contributions that make the work stand out.

MotionBlind
video large language models
motion understanding
contrastive benchmark
world model
🔎 Similar Papers
D
Dhairya Bhatia
Electrical and Computer Engineering Department, Northeastern University, Boston, MA, USA
Bishoy Galoaa
Bishoy Galoaa
Northeastern University
Machine Learning
O
Oliver Fritsche
Electrical and Computer Engineering Department, Northeastern University, Boston, MA, USA
S
Shahid Kamal
Electrical and Computer Engineering Department, Northeastern University, Boston, MA, USA
M
Muhammad Obaidullah Abdul Salam
Electrical and Computer Engineering Department, Northeastern University, Boston, MA, USA
U
Umer Saleem
Electrical and Computer Engineering Department, Northeastern University, Boston, MA, USA
O
Om Rastogi
Electrical and Computer Engineering Department, Northeastern University, Boston, MA, USA
F
Frania Felix Chettiar
Electrical and Computer Engineering Department, Northeastern University, Boston, MA, USA
Nesli Erdogmus
Nesli Erdogmus
Electrical and Computer Engineering Department, Northeastern University, Boston, MA, USA
Sarah Ostadabbas
Sarah Ostadabbas
Electrical & Computer Engineering, Northeastern University
Computer VisionMachine LearningArtificial IntelligenceAugmented Cognition with Medical