LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents

📅 2026-06-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of multi-hop reasoning across tens of thousands of pages in nuclear regulatory document review, where evidence is highly dispersed. The authors propose a state-aware planning framework based on large language models (LLMs) that formulates reasoning as dynamic navigation within an unvectorized document tree. An agent progressively constructs and updates an internal dynamic knowledge graph through browsing, reading, and search actions until sufficient evidence is gathered. A novel auditable edge-reasoning module is introduced to enhance decision traceability without requiring offline indexing. Evaluated on the NuScale FSAR 200-question benchmark, the method achieves 81.5% accuracy and a RAGAS faithfulness score of 0.93, substantially outperforming existing approaches such as PageIndex (+38.0 percentage points), LightRAG, HippoRAG, and GraphRAG.
📝 Abstract
Reviewing nuclear regulatory documents requires multi-hop reasoning across tens of thousands of pages, where judgments depend on evidence assembled across multiple chapters. We frame this task as planning: an LLM-based agent observes the evidence collected so far, picks the next document fragment to inspect, and stops when the evidence is sufficient. The agent operates over a vectorless document tree using browse, read, and search tools, and maintains a dynamic knowledge graph (KG) as state. On a 200-question benchmark over NuScale Final Safety Analysis Report (FSAR) documents, the system reaches 81.5% accuracy with a RAGAS Faithfulness of 0.93. The dominant performance factor is planning: against PageIndex, which uses the same document tree without state-conditioned action selection, the gap is +38.0pp (43.5% to 81.5%, p<0.001). The system also outperforms LightRAG (73.0%, p<0.05), HippoRAG (70.5%, p<0.01), and GraphRAG (49.5%, p<0.001), and matches RAPTOR (75.5%, p=0.11) without offline indexing. Edge inference adds 2.8x cost without raising accuracy; we retain it as a traceability module. Of 7,391 inferred edges, 3 Violates edges (0.04%) flag scope boundaries (Q058) and partial conformance (Q176) as typed annotations that a human reviewer can audit.
Problem

Research questions and friction points this paper is trying to address.

multi-hop reasoning
nuclear regulatory documents
evidence assembly
document review
multimodal reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-guided planning
multi-hop reasoning
dynamic knowledge graph
nuclear regulatory documents
state-conditioned action selection
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Mingyu Jeon
MODULABS, Seoul, Republic of Korea
B
Bokyeong Kim
Korea Atomic Energy Research Institute, Daejeon, Republic of Korea; University of Science & Technology, Daejeon, Republic of Korea
S
Suwan Cho
MODULABS, Seoul, Republic of Korea
J
Jae Young Suh
MODULABS, Seoul, Republic of Korea
Yonggyun Yu
Yonggyun Yu
Principal Researcher, AI Lab, Korea Atomic Energy Research Institute
Deep learningTopology optimizationMusical Instrument