What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience

📅 2026-05-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the systematic discrepancies between synthetic user behaviors generated by large language models (LLMs) and real human actions, which threaten the validity of user experience (UX) research. Conducting first-click tests across 12 authentic UX tasks with 3,431 participants, this work presents the first quantitative evaluation of GPT’s ability to predict both human click behavior and underlying reasoning processes. Synthetic responses were generated using role prompting, chain-of-thought instructions, and varied sampling parameters, then compared against empirical click data. Results reveal that in 53% of tasks, GPT’s predicted distributions significantly diverge from actual human behavior. The findings indicate that cognitive biases stemming from LLMs’ inherent statistical properties are not easily mitigated through prompt engineering; current optimization strategies enhance surface-level plausibility without improving behavioral fidelity, thereby highlighting critical limitations in deploying such models for UX decision-making.
📝 Abstract
Synthetic participants represent a methodologically concerning concept that threatens the integrity of UX research. Findings from previous experiments specify how AI outputs are misaligned with the behaviors and thoughts of real humans in various ways. However, industry voices keep underestimating their severity, advocating for practical compromises where good-enough data does not need to be perfect, and all issues will be solved by future tuning. Our study tackles the lack of systematic understanding of the practical issues that arise with synthetic behavior and its use for steering decisions within real contexts. Within twelve diverse first click tests (n = 3431) obtained from real UX practice, we examine the ability of GPT to predict where humans click and how they reason about their behavior. Results (e.g., significantly different distribution from real data in 53% of tasks) demonstrate critical failures to reflect the patterns in which users click on visual elements and the underlying cognitive processes. Participant personas, chain-of-thought reasoning in GPT, and different sampling parameters fail to create sensible fidelity improvements apart from inflating believability. We expose a multitude of nuanced distortions in synthetic responses that reduce their overall analytical usefulness as a decision-making resource, compared with real data. Observed distortions can be theoretically linked to the properties categorically inherent to LLMs: their statistical nature and encoding of semantic heuristics dependent on their training on linguistic data.
Problem

Research questions and friction points this paper is trying to address.

synthetic participants
human-AI behavioral misalignment
user experience
first click tests
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

synthetic participants
behavioral misalignment
first click test
LLM limitations
UX research
E
Eduard Kuric
Faculty of Informatics and Information Technologies, Slovak University of Technology, Ilkovicova 2, Bratislava, 84216, Slovakia; UXtweak Research, UXtweak j.s.a., Cajakova 18, Bratislava, 81105, Slovakia
P
Peter Demcak
UXtweak Research, UXtweak j.s.a., Cajakova 18, Bratislava, 81105, Slovakia
M
Matus Krajcovic
Faculty of Informatics and Information Technologies, Slovak University of Technology, Ilkovicova 2, Bratislava, 84216, Slovakia; UXtweak Research, UXtweak j.s.a., Cajakova 18, Bratislava, 81105, Slovakia