๐ค AI Summary
This study investigates the alignment between large language models (LLMs) simulating user visual design preferences and actual human preferences, revealing systematic biases and risks of misguidance in AI-assisted design. Drawing on 29 real-world user tests from the UXtweak platform (n=2073), the authors conduct multimodal simulation experiments by manipulating LLM reasoning strategies, sampling approaches, role assignments, and prompt specificity. For the first time in authentic design contexts, they systematically demonstrate that LLM-generated simulations exhibit significant and consistent deviations from real user judgments, producing explanations that lack genuine perceptual grounding and semantic depth while over-relying on superficial, generalized features. This work establishes an empirical foundation and methodological framework for evaluating the algorithmic fidelity of LLMs in humanโcomputer interaction scenarios.
๐ Abstract
Designers of digital solutions increasingly consult Large Language Models (LLMs) for their work. However, it remains unclear how this may affect the user experiences they produce and there are no established practices. We investigate how design preferences expressed by LLM-driven simulation methods align with those of real users. We present a study that aggregates real-world data and design stimuli from twenty-nine preference tests conducted in practice by users of the UXtweak online research platform (n = 2073). We perform holistic multimodal simulations where we manipulate LLM variables (model reasoning, sampling, persona type, and specificity) and assess their effects on algorithmic fidelity. Our results unveil significant and systematic discrepancies between peoples' real design preferences and LLM simulations that are consistent across manipulations. Synthetic justifications lack genuine depth, nuance and reasoning, which they substitute by patterns like focus on generic properties, specific elements, elaboration and overpraising. The unique attention directed by this research toward preferences within visual design stimuli highlights misrepresentation of perception and meaning by LLMs in a context that is intuitive yet critical for design teams. The external and ecological validity of our findings is high, given their replication across a multitude of real-world studies.