LLM-Derived Preference Judgments Are Not Self-Consistent

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过统计测试和可解释度量方法,揭示了基于大型语言模型的偏好判断在自我一致性上的显著不一致问题。
📝 Abstract
Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work estimates a utility function from these judgments and then chooses actions based on their estimated utility. This pipeline assumes the judgments are approximately self-consistent: that a single utility function can reproduce them. But are they? To study this question, we measure the self-consistency of cardinal LLM preference judgments. For example, the difference in stated willingness-to-pay between two items should match the stated payment that makes a person indifferent to exchanging them. We develop statistical tests and interpretable measures of how far observed responses depart from the best-fitting self-consistent utility function. Experiments with flight, apartment, and hotel examples across six LLMs reveal large persistent inconsistencies. This suggests that LLM-derived preference judgments cannot be faithfully summarized by a single utility function.
Problem

Research questions and friction points this paper is trying to address.

self-consistency
preference judgments
utility function
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-consistency
preference judgments
utility function
statistical tests
💼 Related Jobs
No related jobs found.
M
Matthew T. Ford
Cornell University
F
Francis Bahk
Cornell University
Jingjing Wang
Jingjing Wang
Professor, School of Cyber Science and Technology, Beihang University
AI for WirelessUAV NetworksSpace-Air-Ground-Sea NetworksCommunication Security
A
Adam S. Jovine
Cornell University
T
Tinghan Ye
Georgia Institute of Technology
D
David B. Shmoys
Cornell University
P
Peter I. Frazier
Cornell University