Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans
This study investigates whether multimodal large language models (MLLMs) exhibit human-like visual search under foveated input. Utilizing the COCO-Search18 dataset, we compared human and model gaze trajectories through gaze-contingent foveal simulation and a triaxial evaluation framework. Results indicate that while MLLMs achieve detection efficiency comparable to or exceeding human performance, their fixation patterns are characterized by low entropy and non-sequentiality, lacking the temporal dynamics inherent to human vision. This work reveals how "outcome alignment" can obscure underlying "process heterogeneity," highlighting critical blind spots in conventional metrics for assessing human-like temporal mechanisms. These findings provide essential empirical evidence to guide next-generation research toward achieving genuine cognitive alignment in artificial visual systems.