Probing the Structure and Dynamics of LLM Value Expression through Value Conflicts

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入冲突驱动的价值探测方法,探讨了大型语言模型价值表达的结构和动态特性,揭示了表达二元性、功能可引导性和有限可塑性三种模式。
📝 Abstract
Ethical evaluation of Large Language Models (LLMs) often characterizes model values as static and monolithic. In contrast, we argue that LLM value expression is better understood as a structured yet dynamic phenomenon. To investigate this, we introduce Conflict-driven Value Probing, a controlled framework that places LLMs in value conflicts and implements four types of interventions that perturb these conflicts to probe LLM value expression. Applying this framework to ten LLMs, we identify three recurring patterns. (1) Expression duality: models shift from broad idealistic orientations in abstract assessment toward more pragmatic priorities in concrete conflicts. (2) Functional steerability: models readily reconfigure their expressed value profiles toward task-defined value objectives. (3) Bounded plasticity: such reconfiguration is not without constraints, i.e. pressure induces a security- and goal-oriented priority shift while negative framing distinguishes protected values from those more amenable to redirection. Together, these findings characterize both the structure and dynamics of LLM value expression: context flexibly reconfigures expressed priorities, yet within behavioral boundaries. This behavioral account provides a foundation for understanding controllability, alignment, and safety in LLMs. Code and data are available at https://github.com/ZeroGen-Lab/CFProbe.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Value Expression
Conflict-driven Value Probing
Ethical Evaluation
Behavioral Boundaries
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conflict-driven Value Probing
Expression duality
Functional steerability
Bounded plasticity
🔎 Similar Papers
No similar papers found.