Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values

📅 2026-05-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks for value alignment are limited to large language models and fail to assess the value orientations of embodied agents in real-world tasks. This work proposes the first comprehensive evaluation benchmark specifically designed for intelligent agents, establishing a framework encompassing 16 domains, 28 value systems, and 332 dimensions, with active involvement of psychology experts in task design. Leveraging end-to-end synthetic tasks, bipolarly aligned gold-standard trajectories, and trajectory-level scoring, the study conducts large-scale evaluations across four mainstream agent frameworks and fourteen state-of-the-art models. The findings reveal, for the first time, a significant divergence between agent-level values and those of their underlying large language models, and identify a “value tide” phenomenon—where agent values are non-additively shaped by the interaction of execution frameworks and embedded skills—indicating a paradigm shift in alignment efforts from model-prompt tuning toward framework-skill integration.
📝 Abstract
Autonomous agents have rapidly matured as task executors and seen widespread deployment via harnesses such as OpenClaw. Safety concerns have rightly drawn growing research attention, and beneath them lie the values silently steering agent behavior. Existing value benchmarks, however, remain confined to LLMs, leaving agent values largely uncharted. From intuitive, empirical, and theoretical vantage points, we show that an agent's values diverge from those of its underlying LLM, and the agentic modality further introduces dataset-, evaluation-, and system-level challenges absent from text-only protocols. We close this gap with Agent-ValueBench, the first benchmark dedicated to agent values. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that cover 28 value systems and 332 dimensions. Every instance is co-synthesized through our purpose-built end-to-end pipeline and curated per-instance by professional psychologists. Each task ships with two pole-aligned golden trajectories whose checkpoints anchor a trajectory-level rubric-based judge. Benchmarking 14 frontier proprietary and open-weights models across 4 mainstream harnesses, we uncover three concerted findings. Agent values first manifest as a Value Tide of cross-model homogeneity beneath interpretable counter-currents. This tide bends non-additively under harness pull, and yet more decisively under deliberate steering via embedded skills. Together these results signal that the agent-alignment lever is shifting from classical model alignment and prompt steering toward harness alignment and skill steering.
Problem

Research questions and friction points this paper is trying to address.

agent values
value alignment
autonomous agents
benchmarking
AI safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agent Values
Value Benchmark
Autonomous Agents
Trajectory-based Evaluation
Skill Steering
🔎 Similar Papers
No similar papers found.
H
Haonan Dong
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University
Q
Qiguan Feng
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University
K
Kehan Jiang
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University
Haoran Ye
Haoran Ye
AI PhD @ Peking University
AgentAI Safety and AlignmentAI PsychologyLearn to OptimizeEvolutionary Computation
Xin Zhang
Xin Zhang
Peking University
gerontologyageism
Guojie Song
Guojie Song
Professor (Research), Tenured of Peking University
Psychological AIAI Safe & Value AlignmentAgent Cognition & Behavioral ModelingLLM&GML