Implicit Personality Representations in Humans and LLMs

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过对比人类和大语言模型Qwen 2.5-7B-Instruct对人格特质的内在表示,验证了两者在人格结构上的一致性,使用了数百万众包的人格评分数据。
📝 Abstract
A century of psychology has found that the trait words people use to describe one another vary, but the relational structure among those traits, which ones go together and which oppose, is strikingly consistent across raters and cultures. We test whether the LLM (Qwen 2.5-7B-Instruct) reproduces this structure in its internal trait representations. From millions of crowd-sourced personality ratings of fictional characters, we build a human implicit-personality matrix over hundreds of traits; from contrastive model activations, we build a matching matrix over the same traits. The two relational structures align strongly (Mantel r = 0.77), and the agreement holds trait by trait as well as in aggregate. Two dominant axes of the model's trait representations recover the social and intellectual dimensions long known to organize human personality impressions, social warmth and intellectual competence. On held-out dialogue, projecting model activations onto these directions yields personality profiles that agree with human ratings. This work establishes a framework that enables comprehensive, human-grounded comparison between internal model trait geometry and the shared structure of human personality impressions.
Problem

Research questions and friction points this paper is trying to address.

Implicit Personality
LLMs
trait representations
human personality impressions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Implicit Personality Representations
Large Language Models
Human Personality Impressions
Trait Geometry