When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the overconfidence exhibited by LLM agents when facing irresolvable conflicts within personal memory. To this end, we introduce TANGLE, the first benchmark specifically designed to evaluate the handling of authentic memory conflicts, alongside CAAP, a conflict-aware action policy. CAAP enables agents to recognize uncertainty, preserve contradictory evidence, and execute appropriate actions rather than forcing singular conclusions. Experimental results demonstrate that while end-to-end memory extraction struggles to retain conflicting relationships, CAAP significantly outperforms fixed-rule strategies in managing such conflicts. Consequently, this approach effectively enhances both the robustness and honesty of LLM agents when confronted with contradictory memory information, offering a principled solution to mitigate hallucinations arising from inconsistent personal knowledge bases.
📝 Abstract
LLM agents increasingly maintain personal memory across sessions, but it can conflict. Preferences depend on context, behavior evolves, and sources can conflict. When a query lacks context, time, or source authority to interpret conflict, treating one memory as definitive converts unresolved conflict into an unjustified, overconfident action. Existing benchmarks recover one answer from conflicting evidence, overlooking whether agents recognize underdetermination, preserve alternatives, seek missing information, and choose appropriate actions. We introduce \underline{T}esting \underline{A}gents' \underline{N}avigation of \underline{G}enuine, \underline{L}atent, and \underline{E}ntangled Memory Conflicts (\textsc{TANGLE}), a benchmark for genuinely unresolvable memory conflicts. It comprises 541 instances across 40 personas and three types: Context-Partitioned Conflict (CPC), Behavior-Oscillation Conflict (BOC), and Source-Contradiction Conflict (SCC). We evaluate two tracks---an oracle track with curated memory and a pipeline track that extracts memory from multi-session dialogues---on five dimensions: conflict perception, causal reasoning, confidence calibration, clarification seeking, and memory faithfulness. Experiments reveal pipeline challenges. With curated memory, models recognize conflicts more reliably than they calibrate actions or seek targeted clarification. With end-to-end pipeline memory, extraction fails to preserve conflict-bearing relations needed for downstream reasoning. Policy comparisons show fixed rules are insufficient when actions must reflect conflict. These findings motivate Conflict-Aware Action Policy (CAAP), which adapts actions to each conflict using available evidence. \textsc{TANGLE} frames conflict handling as recognizing underdetermination, retaining conflicting evidence, and acting without forcing a definitive answer.
Problem

Research questions and friction points this paper is trying to address.

LLM Agents
Personal Memory
Irreducible Conflict
Underdetermination
Benchmark Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

TANGLE Benchmark
Irreducible Memory Conflict
Conflict-Aware Action Policy
Underdetermination Recognition
Confidence Calibration
💼 Related Jobs
No related jobs found.
L
Lu Yang
Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua University
Shusheng Xu
Shusheng Xu
IIIS, Tsinghua University
Reinforcement learningNLPData mining
Z
Zhuoran Li
Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua University
T
Tongkai Yang
Ant Group
Longbo Huang
Longbo Huang
Professor, IIIS, Tsinghua University, ACM Distinguished Scientist
Reinforcement Learning (RL)Deep RLMachine LearningStochastic NetworksPerformance Evaluation