Iron: Intent-Aligned and Retrospective Dual Learning Framework for Enhancing Generalist Virtual Agents

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Iron框架,通过步进循环一致奖励和后见重生产机制解决数据标注成本高、动作意图不精准及失败轨迹利用低效问题,提升跨环境任务性能。
📝 Abstract
Achieving virtual agents capable of automating tasks across diverse digital environments remains a pivotal challenge in Embodied AI. While Multimodal Large Language Models (MLLMs) offer enhanced visual perception and reasoning, their agentic deployment faces three challenges: costly data annotation, imprecise action-intent alignment, and inefficient exploration from discarded failed trajectories. To address these, we introduce Iron, an intent-aligned, self-improved, and annotation-efficient framework for training GUI agents. Iron employs a novel dual learning strategy that utilizes a stepwise cycle-consistent (SCC) reward to achieve fine-grained alignment between low-level actions and high-level intents, thereby improving instruction grounding and intent understanding. Concurrently, Iron introduces a hindsight reproduction mechanism to repurpose failed trajectories for training, improving both learning efficiency and task diversity. Extensive experiments demonstrate that Iron-trained generalist agents consistently improve performance on cross-environment and cross-device tasks, outperforming models trained with three times more data. Iron also achieves a substantial 25.06% relative improvement on unseen web tasks, with further gains observed on inherently complex tasks, demonstrating the feasibility of building more capable virtual agents.
Problem

Research questions and friction points this paper is trying to address.

Embodied AI
Multimodal Large Language Models
data annotation
action-intent alignment
exploration efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

dual learning strategy
stepwise cycle-consistent reward
hindsight reproduction mechanism
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jiahe Ying
Fudan University
W
Wendong Bu
Zhejiang University
Kaihang Pan
Kaihang Pan
Zhejiang University
nlpvision-and-language
B
Bingchen Miao
Zhejiang University
S
Siyu Chen
Zhejiang University
W
Wen Wang
Zhejiang University
X
Xueming Jiang
Zhejiang University
Juncheng Li
Juncheng Li
East China Normal University
Super ResolutionImage RestorationComputer VisionMedical Image Analysis
Siliang Tang
Siliang Tang
Professor of Computer Science, Zhejiang University
Natural Language ProcessingCross-media AnalysisGraph Neural Network