🤖 AI Summary
This study addresses the systematic discrepancies between synthetic user behaviors generated by large language models (LLMs) and real human actions, which threaten the validity of user experience (UX) research. Conducting first-click tests across 12 authentic UX tasks with 3,431 participants, this work presents the first quantitative evaluation of GPT’s ability to predict both human click behavior and underlying reasoning processes. Synthetic responses were generated using role prompting, chain-of-thought instructions, and varied sampling parameters, then compared against empirical click data. Results reveal that in 53% of tasks, GPT’s predicted distributions significantly diverge from actual human behavior. The findings indicate that cognitive biases stemming from LLMs’ inherent statistical properties are not easily mitigated through prompt engineering; current optimization strategies enhance surface-level plausibility without improving behavioral fidelity, thereby highlighting critical limitations in deploying such models for UX decision-making.
📝 Abstract
Synthetic participants represent a methodologically concerning concept that threatens the integrity of UX research. Findings from previous experiments specify how AI outputs are misaligned with the behaviors and thoughts of real humans in various ways. However, industry voices keep underestimating their severity, advocating for practical compromises where good-enough data does not need to be perfect, and all issues will be solved by future tuning. Our study tackles the lack of systematic understanding of the practical issues that arise with synthetic behavior and its use for steering decisions within real contexts. Within twelve diverse first click tests (n = 3431) obtained from real UX practice, we examine the ability of GPT to predict where humans click and how they reason about their behavior. Results (e.g., significantly different distribution from real data in 53% of tasks) demonstrate critical failures to reflect the patterns in which users click on visual elements and the underlying cognitive processes. Participant personas, chain-of-thought reasoning in GPT, and different sampling parameters fail to create sensible fidelity improvements apart from inflating believability. We expose a multitude of nuanced distortions in synthetic responses that reduce their overall analytical usefulness as a decision-making resource, compared with real data. Observed distortions can be theoretically linked to the properties categorically inherent to LLMs: their statistical nature and encoding of semantic heuristics dependent on their training on linguistic data.