KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出KC-Bench,通过模拟环境和多源冲突任务评估大语言模型在知识冲突、输入不一致及时间冲突中的处理能力。
📝 Abstract
As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts. Its 238 tasks are manually screened from more than 1,000 generated candidates and combine a user simulator, stateful tools, deterministic environment assertions, an open-source natural-language evaluator, and human trajectory verification. Evaluation of nine models, including DeepSeek-V4-Flash, GLM-5.2, and MiniMax-M3, shows substantial cross-domain variation: no model handles factual correction, identity consistency checking, and temporal conflict resolution reliably across all settings. In the simulated environments, missed conflicts can propagate to tool calls or synthetic protected-data flows. KC-Bench isolates this model-level behavior rather than ranking complete agent frameworks, and provides a reproducible diagnostic for developing conflict-aware reasoning and execution safeguards.
Problem

Research questions and friction points this paper is trying to address.

Knowledge Conflicts
LLM Agents
Dynamic Interactive Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic Interactive Benchmark
Knowledge Conflicts
Multi-turn Evaluation
Stateful Tools
Temporal Conflicts
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yaxing Lyu
Shanghai Artificial Intelligence Laboratory; The University of Hong Kong
S
Shengjie Zhou
Xiamen University Malaysia
B
Binbin Toh
Xiamen University Malaysia
Pengyu Zhu
Pengyu Zhu
North China Electric Power University
Artificial IntelligenceBrain-Computer InterfaceAI for SciencePattern Recognition
L
Lijun Li
Shanghai Artificial Intelligence Laboratory