RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts

πŸ“… 2026-07-18
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge that current AI agents struggle to effectively memorize, track, and reason about the evolution of facts across extended interactive contexts, particularly failing to accurately assess the causal impact of fact revisions on downstream conclusions. To this end, the paper introduces RECON, a novel benchmark comprising 24 cross-domain narrative documents (each 50k–100k words), which for the first time focuses on causal tracing and counterfactual reasoning under factual changes. RECON systematically evaluates agents’ compositional reasoning capabilities across six task types, including reconstructing multi-hop evidence chains, handling cascading failures, and resolving source conflicts. Using an integrated retrieval-and-reasoning evaluation framework, experiments reveal that even the strongest non-oracle systems achieve only 22.4% accuracy, highlighting significant limitations of existing approaches in scenarios involving dynamic knowledge evolution.
πŸ“ Abstract
Large language models and LLM-based agents are widely used as personal chat assistants, enterprise copilots, and autonomous workflow agents. In all these applications, memory (the ability to retain, access, and reason over information accumulated over long contexts and multiple interactions) plays a crucial role in determining the reliability of any agent. We introduce RECON (Reasoning over Extended Contexts with Obfuscated Narratives), a benchmark for evaluating compositional reasoning over long contexts. RECON spans 24 case files across three domains (criminal, medical, and financial), each ranging from 50k to 100k tokens, and tests agents on six memory intensive tasks: reconstructing multi-hop evidence chains, propagating cascading invalidations, resolving source conflicts, counterfactual reasoning, satisfying temporal constraints, and temporal fact retrieval. Recent memory benchmarks evaluate whether agents can retrieve scattered facts or detect if a fact has changed whereas RECON evaluates what happens after the change, whether agents can trace which downstream conclusions are affected, which survive through independent support, and how alternative timelines would have unfolded. Our evaluation reveals substantial limitations across current architectures: even the strongest non-Oracle system reaches only 22.4% Accuracy, with retrieval and reasoning each surfacing as challenges.
Problem

Research questions and friction points this paper is trying to address.

agent memory
compositional reasoning
long-context reasoning
memory benchmark
temporal reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

compositional reasoning
long-context memory
agent benchmarking
temporal reasoning
cascading invalidation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
M
Mihir Shriniwas Arya
Department of Computer Science and Engineering, RV College of Engineering