NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the confounding of numerical and visual factors in existing counting benchmarks for Vision-Language Models (VLMs) by constructing a cognitively inspired diagnostic benchmark that enables decoupled evaluation through orthogonal manipulation of visual variables. Employing synthetic data generation, zero-shot testing, and hierarchical probing, the research reveals that model architecture accounts for 32.5% of performance variance, with disparities primarily attributable to the language model component. Furthermore, it demonstrates that early vision encoders already exhibit linearly separable representations of numerosity signals. By elucidating the intrinsic mechanisms and bottlenecks underlying VLM number perception, this work provides a theoretical foundation for enhancing numerical reasoning capabilities in multimodal systems.
📝 Abstract
Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability that emerges in human infants before language acquisition, remains poorly understood in current models, as existing counting benchmarks entangle numerosity with correlated visual factors. We introduce a cognitively inspired diagnostic benchmark, NumerosityVLM, comprising 10,800 synthetic images across six controlled conditions. The benchmark orthogonally manipulates object size, spatial arrangement, and numerosity, while progressively ablating texture, shape, and color. Evaluating seven VLMs in a zero-shot setting, multi-factor analysis reveals that model architecture explains the largest proportion of performance variance (partial $ω^{2}=0.325$), far exceeding visual conditions. Layer-wise probing further shows that linearly separable numerosity signals consistently emerge at early stages of the vision encoder, while performance differences across evaluated models are primarily associated with the language model component. Code and data are publicly available at https://github.com/fuy3/NumerosityVLM-Benchmark, and https://huggingface.co/datasets/fuy3/NumerosityVLM.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Numerosity Perception
Counting Benchmarks
Visual Confounds
Innovation

Methods, ideas, or system contributions that make the work stand out.

NumerosityVLM
Cognitively Inspired Benchmark
Numerosity Perception
Orthogonal Manipulation
Layer-wise Probing
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Yiming Fu
Yiming Fu
HKUST
bio-photonicsadaptive opticsneuroscience
Fangjun Li
Fangjun Li
Cognitive Robotics Lab, Department of Computer Science, The University of Manchester, Manchester M13 9PL, UK
X
Xiujin Liu
University of Michigan, MI 48109, USA
R
Ruidong Ma
Sheffield Hallam University, Sheffield S1 1WB, UK
H
Hang Yu
Tufts University, MA 02155, USA
Z
Zhichen Lu
ENSTA, Institut Polytechnique de Paris, Palaiseau 91120, France
K
Kanwei He
Cognitive Robotics Lab, Department of Computer Science, The University of Manchester, Manchester M13 9PL, UK
Alessandro Di Nuovo
Alessandro Di Nuovo
Professor of Machine Intelligence, Sheffield Hallam University (UK)
Computational IntelligenceCognitive SystemsRoboticsComputer Science
Angelo Cangelosi
Angelo Cangelosi
Professor of Machine Learning and Robotics, University of Manchester
Artificial intelligenceroboticsmachine learningcognitive science
Z
Zhegong Shangguan
Cognitive Robotics Lab, Department of Computer Science, The University of Manchester, Manchester M13 9PL, UK