GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing Vision-Language-Action (VLA) models in heterogeneous scaling and cross-embodiment generalization by proposing a tri-system embodied foundation model. By unifying cognitive prediction with action generation, the framework leverages 37,000 hours of heterogeneous data for pre-training alongside a single-stage alignment strategy to jointly optimize comprehension and multi-embodiment control. Experimental results demonstrate significant improvements in zero-shot reasoning and instruction-following capabilities. Furthermore, the model exhibits superior task adaptability and completion rates across both domestic and industrial environments, effectively validating the efficacy of the proposed architecture and the critical value of large-scale heterogeneous data in advancing embodied intelligence.
📝 Abstract
Vision-language-action (VLA) models have become a dominant paradigm for generalist embodied agents, demonstrating strong complex and long-horizon task completion in structured settings. Yet it remains an open question whether current VLA systems can benefit from more effective architectural design, scale to substantially larger and more heterogeneous data regimes, and achieve broader generalization across tasks and embodiments. To this end, we present GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments. Specifically, GigaBrain-0.7 unifies understanding, prediction, and action through a three-system architecture, scales pretraining to over 37,000 hours of heterogeneous embodied data, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the preceding GigaBrain-0 series and prior state-of-the-art models including $π_{0.5}$, GigaBrain-0.7 achieves substantial improvements in foundation zero-shot capabilities, language-conditioned instruction following, and post-training task success rates. In particular, on our in-house Maker H01 platform and mainstream robot embodiments, GigaBrain-0.7 demonstrates strong task adaptability and completion ability across both home and industrial scenarios. All training code and pretrained model weights will be released.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Embodied Foundation Models
Cross-embodiment Generalization
Heterogeneous Data Scaling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Three-System Architecture
One-Stage Alignment Training
Heterogeneous Embodied Data
Embodied Foundation Model
Zero-Shot Generalization
🔎 Similar Papers
No similar papers found.