🤖 AI Summary
This study addresses the lack of systematic evaluation regarding whether existing vision models adhere to human Gestalt principles in visual organization. The authors introduce the first reusable benchmark for Gestalt perception, constructing a behavioral test suite based on four canonical Gestalt grouping tasks and leveraging published human perceptual data to evaluate 45 prominent vision models—including supervised, self-supervised, contrastive vision-language encoders, and both open- and closed-source foundation models—without requiring new human experiments. Their analysis reveals that certain high-accuracy closed-source models exhibit significant deviations from human-like Gestalt consistency, highlighting perceptual biases that conventional performance metrics fail to capture.
📝 Abstract
Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects. These are the Gestalt operations that visualization design builds on. Whether vision models organize visual content this way has not been systematically tested. We introduce a behavioral battery that scores models against human data from prior perception studies on four grouping tasks: mark-color odd-one-out, color-series counting, silhouette recognition, and object odd-one-out. We apply it to 45 models across five training families: supervised, self-supervised, and contrastive vision-language encoders, open-weight VLMs, and closed foundation models. The battery reveals that agreement with human responses captures aspects of perceptual organization that conventional performance metrics fail to distinguish, with several closed models exhibiting substantially lower alignment than their benchmark accuracy would suggest. Scoring against published perception data therefore gives visualization research a reusable yardstick, requiring no new user study, for auditing whether the models now entering visualization pipelines organize what they see the way their human audience does.