LexKairos: Benchmarking Legal Temporal Capabilities in LLMs

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the absence of systematic evaluation of large language models’ (LLMs) temporal reasoning capabilities in legal contexts—a critical gap given the centrality of time-related elements in legal practice. The authors propose LexKairos, the first multidimensional benchmark tailored to Chinese legal scenarios, encompassing three core dimensions: temporal knowledge in statutes, temporal modeling in case analysis, and statute-case temporal reasoning. It comprises nine subtasks grounded in real judicial cases and statutory provisions. The study systematically evaluates eight prominent LLMs under diverse reasoning paradigms, including Vanilla prompting, Chain-of-Thought, and Thinking protocols. Results indicate that Gemini-3-Flash achieves the strongest overall performance; however, all models exhibit significant deficiencies in tasks requiring precise recall of temporal metadata or complex deadline-based reasoning, thereby exposing fundamental challenges in current LLMs’ capacity for legal temporal understanding.
📝 Abstract
Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that governs the validity of statutes, the progression of legal cases, and the enforcement of procedural deadlines. However, legal temporal capabilities remain underexplored in existing legal AI benchmarks. To address this gap, we propose LexKairos, a comprehensive benchmark for evaluating the temporal capabilities of LLMs in the Chinese legal context across three dimensions: statutory temporal knowledge, case temporal modeling, and statute-case temporal reasoning. LexKairos comprises nine sub-tasks drawn from real-world Chinese judicial cases and statutes. We conduct systematic evaluations of eight LLMs under multiple inference settings, including vanilla, Chain-of-Thought (CoT), and thinking modes. Our results show that Gemini-3-Flash achieves the strongest overall performance, yet even the best-performing model exhibits notable limitations on tasks demanding precise time-sensitive statutory metadata recall or complex reasoning in time limits, indicating that legal temporal knowledge and reasoning remain open challenges for current LLMs. Data and code are available at https://github.com/thunlp/LexKairos.
Problem

Research questions and friction points this paper is trying to address.

legal temporal reasoning
large language models
temporal capabilities
legal AI benchmarking
statutory time limits
Innovation

Methods, ideas, or system contributions that make the work stand out.

legal temporal reasoning
LLM benchmarking
statutory time knowledge
case temporal modeling
Chinese legal AI