π€ AI Summary
This study addresses the oversight of dissociated human attention and recognition mechanisms in existing AI image editing detection. We propose a two-stage cognitive model demonstrating that edited regions drive attentional capture while semantic plausibility determines judgment accuracy. As the first work to introduce pre-attentive and recognition distinctions into this domain, we construct a generative eye-movement prediction framework. Validated through eye-tracking and mixed-effects analyses, the significant dissociation between stages is confirmed. The model achieves attention prediction correlations of 0.77β0.82, and its missed-detection behavioral prediction performance (r=0.52) significantly outperforms linear baselines (r=0.48). These findings establish a novel paradigm for understanding detection blind spots in human-AI interaction, highlighting the critical role of cognitive separation in evaluating synthetic imagery.
π Abstract
As AI-generated image edits proliferate, the platforms meant to curb the resulting disinformation treat detectability as a single, undifferentiated property: an edit either gets a warning or it does not. We show this is the wrong model. Across a controlled eye-tracking study ($N=59$, Latin-square design, four conditions crossing edit area and semantic plausibility), a mixed-effects analysis reveals that whether an edit is noticed and whether it is correctly judged as fake are dissociable stages, governed by different factors: edit area drives attention capture ($p<0.001$) while semantic plausibility drives judgment accuracy and look-but-fail-to-see (LBFS) error rates ($p<0.001$). This dissociation survives correction for multiple comparisons; a secondary interaction between the two factors does not. This two-stage account extends a long-standing distinction in visual attention research (between pre-attentive capture and effortful recognition) into the new domain of AI-edit detectability. We then test whether a generative eye-movement model can computationally operationalize the attention-capture stage: a Transformer trained to generate scanpaths tracks per-image attention with strong discriminative power (Pearson $r=0.77$--$0.82$ across held-out stimuli) and, on the harder task of predicting LBFS incidence, modestly outperforms a two-parameter linear baseline even without access to the plausibility label ($r=0.52$ vs. $r=0.48$). We report this comparison, our ablations, and our method's limitations (a single fixed train/validation split, not leave-one-subject-out) without inflation, consistent with responsibly communicating what a machine learning system can and cannot do to help curb AI-driven disinformation.