Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入无模型的渲染天花板方法,解决了视觉-语言模型中感知与推理分离的问题,有效评估了模型在晶体结构上的表现。
📝 Abstract
Multimodal evaluations cannot say whether a vision-language model misread an image or misreasoned about it, because every existing method for separating the two places a second model in the loop. We introduce the render ceiling, a model-free reference for benchmarks built by rendering known objects: inverting the frozen cameras and re-solving cross-view correspondence recovers exactly the answer the images support. We prove the ceiling fails only through an enumerable set of projection coincidences and certify that set empty on 2,160 rendered crystal structures, so every point of a model's deficit belongs to the model. Across fourteen vision-language models, supplying exact geometry as text lifts every model yet closes under half the gap for thirteen, while a supervised vision model with no language component reads the same images at 0.8952, above every vision-language model. The instrument exposes extraction-stage fabrication that downstream accuracy would misattribute to reasoning, yields camera-placement rules for benchmark builders, and transfers to any benchmark with an invertible forward rendering.
Problem

Research questions and friction points this paper is trying to address.

vision-language model
perception
reasoning
multimodal evaluation
render ceiling
Innovation

Methods, ideas, or system contributions that make the work stand out.

render ceiling
model-free reference
cross-view correspondence
perception and reasoning separation
crystal structures