TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of static temporal question-answering benchmarks by introducing the first dynamic, real-time benchmark spanning six domains and sixty scenarios, alongside a self-evolving agent framework equipped with a reusable skill library. Through multi-granularity data release and an LLM-based evaluation system, this work systematically assesses model capabilities in state recognition and prediction within evolving environments. The results reveal significant performance gaps in frontier models regarding temporal validity and context utilization. Furthermore, by providing monthly updated resources and comprehensive failure mode analyses, this research establishes a new paradigm for investigating dynamic temporal reasoning, overcoming the lack of timeliness assessment in existing methodologies.
📝 Abstract
Time series analysis in high-stakes domains relies on recurring data releases, where new observations can alter the evidence base and the validity of later conclusions. Existing time series QA benchmarks mostly rely on fixed snapshots, leaving temporal validity and cutoff-aware evidence use unevaluated. We introduce TimeSage-EV, a live benchmark for agentic time series analysis in evolving environments. It tracks 60 real institutional scenarios across 6 domains, comprising 1,485 scenario-period QA pairs from Feb 2023 to May 2026 and spanning monthly, weekly, daily, and irregular release cadences. At each period, large language model (LLM) agents receive time series data and source reports, while the withheld target release provides ground truth. TimeSage-EV evaluates state identification, data summarization, and outlook reasoning. Experiments with frontier LLM agents and TimeSage-1.0, a novel self-evolving agent with a reusable analytical skill library, reveal significant performance gaps across model tiers and recurring failures in temporal validity, exogenous context use, and adaptation. We release TimeSage-EV as a research resource with monthly updates, code, a leaderboard, and failure-mode analyses.
Problem

Research questions and friction points this paper is trying to address.

Time Series Analysis
Evolving Environments
Temporal Validity
LLM Agents
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Live Benchmark
Evolving Environments
Self-evolving Agent
Temporal Validity
Reusable Skill Library
🔎 Similar Papers