RigidBench: Evaluating Rigid-Body Physics in Video Generation Models

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of rigid-body physics evaluation metrics in video generation by constructing a simulator-based physical benchmark. We propose an evaluation framework that decouples visual fidelity from physical plausibility, alongside a novel 6-DoF trajectory scoring method. Our analysis reveals a counterintuitive positive correlation between SSIM and 3D trajectory error, further investigated through diffusion transformer probing of internal representations. Experiments demonstrate that no existing model achieves comprehensive superiority. However, fine-tuning with our proposed dataset reduces 3D trajectory error by 20% while preserving SSIM, effectively validating that independent physical assessment and targeted optimization are critical for enhancing physical consistency in generated videos.
📝 Abstract
Video models are increasingly used to predict what happens next in a scene, yet the metrics commonly used to compare their outputs say little about whether the predicted objects move correctly. Motion, geometry, identity, background stability, and visual similarity can fail independently, but whole-frame scores often mix these errors together. We introduce RigidBench, a simulator-grounded benchmark that compares a generated continuation with a reference rollout from the same initial frame and motion description. Its five rigid-body tasks vary objects, materials, viewpoints, and indoor and outdoor scenes, with per-frame masks, depth, 6-DoF trajectories, and contacts available for scoring. We evaluate eight models on the same 100 examples with ten measurements that keep these aspects separate. The resulting rankings depend strongly on what is measured: no model leads on all ten, and across model means, higher SSIM accompanies larger 3D trajectory error (r = 0.89). RigidBench also includes 5,000 training videos with exact simulator state, which we use to fine-tune and analyze Wan 2.2 TI2V-5B. Full fine-tuning reduces 3D trajectory error by about 20% with almost no change in SSIM, while teacher-forced probes and targeted interventions show that object position is represented throughout Wan's diffusion transformer and used by its denoising computation.
Problem

Research questions and friction points this paper is trying to address.

Video Generation Evaluation
Rigid-Body Physics
Motion Correctness
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

RigidBench
Simulator-grounded Evaluation
Rigid-body Physics
Decoupled Metrics
Diffusion Transformer Probing
🔎 Similar Papers
No similar papers found.