PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过将视觉-语言模型转化为物理基础模型,解决了理解物理环境、生成动作和预测未来状态的问题,采用自回归序列优化方法,并结合人类交互视频预训练及监督微调。
📝 Abstract
We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.
Problem

Research questions and friction points this paper is trying to address.

Physical Environments
Action Generation
Future State Prediction
Unified Model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Model
Autoregressive Next-Token Prediction
Embodied Supervision
Multimodal Capabilities
D
DeepCybo Team
Y
Yu Bin
H
Haipeng Cao
Z
Zheng Chang
K
Kai Chen
Y
Youning Chen
K
Kailin Deng
Y
Yichao Du
X
Xiaotong Fu
H
Haoyang Ge
Y
Yunlong Guo
C
Chenliu Hao
Jiyan He
Jiyan He
University of Science and Technology of China
Machine LearningAI for Science
X
Xuguo He
Y
Yakun Hou
K
Kai Hu
Cong Huang
Cong Huang
University of Science and Technology of China
Image/Video processing
T
Tuopusen Huang
Y
Yu Huang
H
Hong Li
P
Peize Li
S
Shijie Lian
X
Xiaopeng Lin
Y
Yun Lin
H
Haibao Liu