GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of physical hallucination, poor generalization, and inadequate environmental perception in large language model (LLM)-driven embodied agents when performing long-horizon tasks. To overcome these limitations, the authors propose a synergistic planning framework that integrates task graphs and scene graphs. The task graph encodes structured prior knowledge to guide high-level planning, while the scene graph enables event-driven dynamic replanning, establishing a closed-loop perception and error-correction mechanism. This approach is the first to incorporate dual-graph structures into LLM-based planning pipelines, combining contextual prompting, GRPO-based reward design, and model fine-tuning. Evaluated on the ALFRED benchmark, the method achieves state-of-the-art performance, significantly improving task success rates under both zero-shot and few-shot settings and demonstrating strong out-of-distribution generalization on unseen long-horizon tasks.
📝 Abstract
Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack of environmental awareness. We propose GraphThink, a novel framework that integrates a task graph to provide structured knowledge for robust planning and a scene graph to maintain environmental memory for event-driven replanning. Specifically, the task graph guides LLM thinking through contextual prompting and iterative refinement, effectively mitigating planning hallucinations. Furthermore, within the GRPO framework, the task graph offers delicate reward design to train the LLM planner, enhancing long-horizon planning capabilities and improving generalization. Finally, an event-driven replanning module, powered by the scene graph, enables closed-loop environment awareness and error correction. GraphThink achieves state-of-the-art performance on the ALFRED benchmark. In particular, our high-level planner surpasses leading API-based LLMs on both the validation set and held-out long-horizon tasks, underscoring its robust zero-shot and few-shot capabilities. Additional evaluations further demonstrate strong out-of-distribution generalization to novel tasks and environments.
Problem

Research questions and friction points this paper is trying to address.

physical hallucinations
long-horizon tasks
environmental awareness
generalization
embodied task planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

GraphThink
task graph
scene graph
long-horizon planning
embodied AI
C
Chen Li
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China; Beijing Key Laboratory of Research on Large Models and Intelligent Governance, Beijing, China
S
Sijie Cheng
Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China; RayNeo.AI, Shenzhen, China
Yuelin Zhang
Yuelin Zhang
Gaoling School of Artificial Intelligence, Renmin University of China
MLLMGenmetric GNN
J
Junxi Li
Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong SAR, China
M
Maozhi Huang
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China; Beijing Key Laboratory of Research on Large Models and Intelligent Governance, Beijing, China
Yang Liu
Yang Liu
Tsinghua University
Wenbing Huang
Wenbing Huang
Associate Professor, Renmin University of China
Machine LearningAI for Science