A Brain-inspired Hierarchical Framework for Zero-Shot Robot Task Reasoning and Execution

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Robots that follow open-ended language instructions need to connect semantic intent to visual scene understanding, geometric feasibility, object states, and physical interaction conditions. End-to-end Vision-Language-Action policies have improved cross-task generalization, but they typically map visual and language inputs directly to robot actions, leaving limited explicit structure for long-horizon decomposition, physical verification, and recovery. We present \method, a zero-shot hierarchical framework functionally inspired by the division of roles in the human brain, comprising visual perception and state inference, language grounding and action-sequence generation from a shared atomic action library, cost-based plan selection, and real-robot execution and verification. The framework grounds commands in explicit object states, composes reusable atomic actions into task-conditioned sequences, ranks alternative sequences by execution cost, and verifies intermediate physical outcomes from refreshed observations. In the evaluation, \method{} completes 10/10 clean board trials, 10/10 pick-and-place trials, and 4/5 pyramid stacking trials for both the flat and irregular initial-layout conditions; the corresponding mean task progress is $99.03\%$, $100.00\%$, and $96.67\%$ respectively. Across all evaluated conditions, \method{} achieves higher success rates than ReKep, Dream2Flow, and $\pi_{0.5}$ benchmarks, demonstrating the effectiveness of combining explicit object-state reasoning, compositional atomic actions, cost-based plan selection, and closed-loop execution verification.
Problem

Research questions and friction points this paper is trying to address.

Zero-Shot Robot Task
Vision-Language-Action
Hierarchical Framework
Semantic Intent
Physical Interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

hierarchical framework
zero-shot task reasoning
cost-based plan selection
closed-loop execution verification
💼 Related Jobs
No related jobs found.
Guangming Wang
Guangming Wang
University of Cambridge, ETH Zurich, and Shanghai Jiao Tong University
Robot VisionRobot ManipulationRoboticsComputer VisionAutonomous Driving
P
Pengfei Ye
Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Cambridge, MA, USA; and Department of Mechanical and Aerospace Engineering, Hong Kong University of Science and Technology, Hong Kong, China
Qizhen Ying
Qizhen Ying
MSc, University of Oxford
Yixiong Jing
Yixiong Jing
University of Cambridge
Point cloudDeep learningGenerative modelsStructural Health monitoring
Y
Yuxiang Ma
Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Cambridge, MA, USA
H
Haonan Chen
Computer Science and Kempner Institute, Harvard University, Cambridge, MA, USA
H
Haibing Wu
Department of Engineering, University of Cambridge, Cambridge, U.K.
Olaf Wysocki
Olaf Wysocki
Assistant Research Professor, University of Cambridge
Computer VisionPhotogrammetryMachine Learning
M
Molong Duan
Department of Mechanical and Aerospace Engineering, Hong Kong University of Science and Technology, Hong Kong, China
Brian Sheil
Brian Sheil
Laing O'Rourke A/Prof, EPSRC Open Fellow @ University of Cambridge; Chief Scientist @ InfraMind
AIcomputer visioninfrastructureconstructiongeomechanics