Discovering Diverse Planning Policies for Multimodal Embodied Agents with Quality-Diversity Optimization

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the persistent stagnation often encountered by multimodal embodied agents in long-horizon tasks due to overreliance on a single planning strategy, which hinders recovery from failure. The study introduces quality-diversity (QD) optimization into embodied planning—proposing a framework that offline constructs a diverse repertoire of strategies and online adaptively switches among them. Diversity is fostered through experience-guided recombination and mutation, while a behavior-space index is built upon interaction intensity and goal-directedness. During execution, the system continuously monitors task progress and, upon detecting stagnation, rolls back and switches to a behaviorally distinct alternative strategy. Evaluated on the ThreeDWorld transportation benchmark, the approach significantly improves task success rates and interaction efficiency, demonstrating that strategic diversity is crucial for adaptive planning and robust failure recovery.
📝 Abstract
Multimodal embodied agents are increasingly required to solve long-horizon tasks by integrating visual observations, textual goals, and interaction history into closed-loop decision making. However, state-of-the-art large-model-based planners often rely on a single dominant planning style during execution. Once this execution mode becomes ineffective, the agent may remain stalled for many steps, repeatedly interacting with the environment without making meaningful progress. We address this limitation by proposing a Quality-Diversity (QD) framework for discovering diverse planning policies for multimodal embodied agents. The proposed method treats planning-policy templates as evolvable individuals and organizes them into a behavior-indexed archive rather than collapsing search to a single prompt style. In the offline stage, rollout trajectories are summarized into structured success and failure experiences, which guide policy variation through recombination and experience-guided mutation. The resulting policies are mapped into a behavior space defined by interaction intensity and goal-directedness, and the highest-quality policy in each niche is retained in the archive. In the online stage, the agent executes one policy at a time while monitoring task progress. When persistent stall is detected, the system rolls back to the latest checkpoint and switches to a behaviorally distinct archive policy to resume execution. Experiments on the ThreeDWorld transport benchmark show that the proposed framework improves both task success and interaction efficiency over representative baseline planners. These results suggest that discovering diverse policy repertoires is an effective way to support adaptive multimodal planning and online failure recovery.
Problem

Research questions and friction points this paper is trying to address.

multimodal embodied agents
long-horizon tasks
planning diversity
execution stall
adaptive planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quality-Diversity Optimization
Multimodal Embodied Agents
Diverse Planning Policies
Behavior Space
Online Failure Recovery
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Pengfei Xu
School of Artificial Intelligence, Nanjing University of Information Science and Technology, Nanjing, Jiangsu, China
Y
Yong Liu
School of Artificial Intelligence, Nanjing University of Information Science and Technology, Nanjing, Jiangsu, China
X
Xiaoya Nan
School of Artificial Intelligence, Nanjing University of Information Science and Technology, Nanjing, Jiangsu, China
Q
Qiang Yang
School of Artificial Intelligence, Nanjing University of Information Science and Technology, Nanjing, Jiangsu, China
P
Peilan Xu
School of Artificial Intelligence, Nanjing University of Information Science and Technology, Nanjing, Jiangsu, China