🤖 AI Summary
This study investigates whether multimodal large language models (MLLMs) exhibit human-like visual search under foveated input. Utilizing the COCO-Search18 dataset, we compared human and model gaze trajectories through gaze-contingent foveal simulation and a triaxial evaluation framework. Results indicate that while MLLMs achieve detection efficiency comparable to or exceeding human performance, their fixation patterns are characterized by low entropy and non-sequentiality, lacking the temporal dynamics inherent to human vision. This work reveals how "outcome alignment" can obscure underlying "process heterogeneity," highlighting critical blind spots in conventional metrics for assessing human-like temporal mechanisms. These findings provide essential empirical evidence to guide next-generation research toward achieving genuine cognitive alignment in artificial visual systems.
📝 Abstract
Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal large language models (MLLMs), given the same foveated input, search as humans do bears on their use as models of human vision and on attention-alignment scores. We compare three general-purpose MLLMs with human eye-movement scanpaths on goal-directed search (COCO-Search18), driving each model fixation by fixation through an identical, human-matched foveated view and assessing it along three axes: the decision of target presence, the efficiency of reaching the target, and the gaze process itself. The axes dissociate. On the decision and on target acquisition the models match or exceed humans, detecting present targets near ceiling and reaching them on the first saccade more often than people do. The gaze process is not human. Under the human-matched condition, all three share one signature: low-entropy, large-amplitude, self-consistent scanpaths that agree with themselves far more closely than two humans agree with each other. That is consistent with a single-pass, non-serial architecture rather than a limit of acuity. Matched retinal input reproduces where humans look but not how the looking unfolds in time, and no degradation regime recovers human-like search at human-like success. The gap sits on a process axis that answer-alignment and saliency metrics do not measure. Because they miss it, such metrics cannot certify human-like vision, and zero-shot models suit outcome and spatial questions but not temporal, process-level ones.