BabyVision: Visual Reasoning Beyond Language
This work addresses the overreliance of current multimodal large language models on linguistic priors and their consequent deficiency in foundational visual understanding—capabilities that even human infants possess—leading to markedly subpar performance on basic visual tasks. To systematically evaluate pure visual reasoning independent of language, the authors introduce BabyVision, a comprehensive benchmark comprising 388 non-linguistic visual tasks across four major categories and 22 subcategories. They further present BabyVision-Gen, a generative model tailored for this benchmark, along with an automated evaluation toolkit. Experimental results reveal that leading models, such as Gemini3-Pro-Preview (scoring 49.7), fall significantly short of adult human performance (94.1), underscoring a critical gap in foundational visual primitives and highlighting the need to advance multimodal models toward more human-like visual perception.