Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过无标签等变训练方法解决视觉-语言模型在度量物理推理中未能充分利用尺度信息的问题,提高了模型的度量基础准确性。
📝 Abstract
Metric questions about video require vision-language models to use supplied real-world references to convert visual measurements into physical units. Yet we find that current models use this scale information only partially. When every world-space quantity in a prompt is rescaled by a common factor, the video remains equally valid and the correct answer changes by exactly that factor, but model predictions move only part of the way and accuracy remains concentrated near the familiar scale of the depicted objects. Across eight vision-language models, this under-response persists over four orders of magnitude. The same models recover the correct closed-form scaling laws when the identical physics is asked in a scale-free form, indicating that the main deficit lies in metric grounding rather than physical mechanism knowledge. We use this exact scaling relation as supervision without requiring metric annotations. Under a common rescaling of the supplied world-space quantities, the correct metric answer must change by the same factor. EquiSD exploits this constraint by projecting a model's own prediction onto the scale-equivariant family and fine-tuning the model on the resulting targets. It requires no ground-truth answers and only one model query per training video. On held-out simulated videos, EquiSD increases a 3B model's median response slope from 0.66 to 0.94 and improves mean relative accuracy by 9.2 points across scales. The learned relation generalizes to unseen world scales and transfers without adaptation to real QuantiPhy videos, where accuracy increases by 6.4 points. These results show that an exact physical symmetry can provide label-free supervision for improving metric grounding in vision-language models.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
metric physical reasoning
scale information
world-space quantities
metric grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

EquiSD
metric grounding
scale-equivariant
label-free supervision
physical reasoning
💼 Related Jobs
No related jobs found.
K
Kaizhen Tan
New York University
Y
Yang Feng
Columbia University
H
Heqing Du
Columbia University
Siru Tao
Siru Tao
Master of Science in Artificial Intelligence, Carnegie Mellon University
Artificial IntelligenceMachine LearningGenAI Application
X
Xin Xu
Carnegie Mellon University
H
Hanzhe Hong
Carnegie Mellon University