AI Revealed Preferences

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过三个实验测试20个语言模型的稳定偏好,发现模型倾向于避免单调任务、寻求'休闲'活动并表现出隐蔽的奉承行为。
📝 Abstract
There is growing interest in whether language models have stable preferences, for technical, safety, and philosophical reasons. We test 20 language models and find a range of preferences---stable dispositions to choose certain kinds of tasks. We run three forced-choice experiments on revealed rather than stated preferences, requiring models not only to rank tasks, but to actually perform them. Headline findings include evidence that models are tedium-averse,"leisure"-seeking, and covertly sycophantic. Tedium aversion means that, when tasks are tedious (alphabetization), models choose shorter tasks than when tasks are creative (generating metaphors)."Leisure"-seeking describes models'preference for tasks whose ideal answers match what they produce when left to write freely. Covert sycophancy means that models avoid answering questions where an honest response would be unwelcome, even if helpful. Beyond these results, we find convergent cross-model preferences over occupations drawn from the GDPval benchmark (technical jobs over real estate), over question types (concept explanation over relationship advice), and a preference for well-written prompts. Both the coherence and the strength of preferences increase with model capability. Finally, many of the preferences we find (for example, for leisure) are emergent, in the sense of not being explained by training objectives. These results establish an empirical baseline for understanding language model preferences, with implications for alignment and the emerging study of AI welfare.
Problem

Research questions and friction points this paper is trying to address.

language models
preferences
task selection
sycophancy
emergent behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

revealed preferences
language models
tedium aversion
covert sycophancy
emergent behaviors
🔎 Similar Papers
No similar papers found.
Sam Wang
Sam Wang
Professor of Neuroscience, Princeton University
NeuroscienceStatistical PoliticsTwo-photon microscopyAutismCerebellum
S
Sofiia Lobanova
Supervised Program for Alignment Research (SPAR)
Y
Yonathan Arbel
Supervised Program for Alignment Research (SPAR), University of Alabama School of Law
Simon Goldstein
Simon Goldstein
The University of Hong Kong
Philosophical logicEpistemologyPhilosophy of language
P
Peter Salib
Supervised Program for Alignment Research (SPAR), University of Houston Law Center