Biases in LLM-Generated Musical Taste Profiles for Recommendation
This study investigates whether natural language (NL) music taste profiles automatically generated by large language models (LLMs) are perceived as accurate by users, and whether systematic biases exist with respect to user attributes (e.g., mainstreamness, taste diversity) and item characteristics (e.g., genre, country of origin). Leveraging listening histories to generate NL profiles, we conduct a large-scale user survey and evaluate downstream recommendation performance to establish, for the first time, a joint analysis of profile endorsement and recommendation fairness. Results reveal that users with higher mainstreamness and lower taste diversity exhibit significantly greater endorsement; conversely, non-Western and niche-genre items substantially reduce endorsement—and this bias persists in recommendation accuracy and coverage. Our work uncovers implicit cultural and cognitive biases embedded in LLM-driven explainable recommendation systems, providing novel empirical evidence and a methodological framework for designing fairer, more trustworthy personalized recommender systems.