Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了大型语言模型(LLMs)的价值观对齐问题,通过量化分析和PEC框架,实现更高效精准的调整,避免了盲目训练。
📝 Abstract
As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictably under minor wording changes ("swing"), yet stubbornly ignore explicit instructions to correct ingrained biases ("rigidity"). Resolving this duality is critical for reliable AI alignment. To systematically understand and safely steer these latent subjective preferences, our study is structured around three fundamental questions. First, do LLMs possess an intrinsic value system? By projecting responses from 106 LLMs (150,000 queries per model) and 95,000 human survey profiles into a shared sociological space, we empirically confirm that they do. However, they do not mirror human diversity, instead crystallizing into a highly concentrated, idealized value core. Second, how can these values be quantified? We propose the Prior-Environment-Cognition (PEC) framework. This model mathematically defines value expression as the joint outcome of inherent dispositions like parameter weights (Prior), external contexts such as user prompts (Environment), and internal reasoning processes like Chain-of-Thought (Cognition). Finally, how can LLMs' values be aligned toward a desired target? Using PEC diagnostics, we establish an adaptive "Alignment Prescription". Rather than blindly applying resource-intensive training, this method identifies the minimum effective intervention needed for each dimension, ranging from zero-cost prompts to targeted parameter updates. Extensive empirical validation confirms that our approach successfully verifies the presence of LLM values, accurately quantifies their shifts, and achieves more efficient and precise steering than conventional blind training, all without degrading general capabilities.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Value Alignment
Behavioral Paradox
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantitative Analysis
Alignment Framework
Prior-Environment-Cognition (PEC)
Adaptive Alignment Prescription
🔎 Similar Papers
K
Keqing Zhang
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, 95 Zhongguancun East Road, Beijing, 100190, China; School of Industry-education Integration, University of Chinese Academy of Sciences, 1 Yanqihu East Road, Huairou, Beijing, 101499, China
Jingyu Chen
Jingyu Chen
Huazhong University of Science and Technology
Computer VisionDeep Learning3D Vision
Yufan Liu
Yufan Liu
Institute of Automation, Chinese Academy of Sciences
Image/video processingKnowledge DistillationSaliency detectionModel compressionVideo coding
Y
Yongqiang Zhu
Beijing Jiaotong University, 3 Shangyuancun, Beijing, 100044, China
Nai Ding
Nai Ding
Zhejiang University
speech perceptionlanguageauditory neuroscience
Lai Jiang
Lai Jiang
Beihang University, University of British Columbia
Computer visionMultimediaMedical imaging
Congyan Lang
Congyan Lang
Beijing Jiaotong University
computer vision
B
Bing Li
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, 95 Zhongguancun East Road, Beijing, 100190, China
W
Weiming Hu
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, 95 Zhongguancun East Road, Beijing, 100190, China