Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Object-Uni模型,通过将物体姿态作为共享几何变量,解决了物体实例空间状态的理解与可控生成问题。
📝 Abstract
Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects in natural language, but they struggle to precisely represent continuous object poses and generate geometrically consistent images under target viewpoints. To mitigate this, we propose \emph{Object-Uni}, a unified model for object-centric spatial understanding and controllable generation. Specifically, we formulate object-centric spatial intelligence as a unified problem connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis. We treat object pose as an explicit geometric variable shared by understanding and generation, rather than merely a prediction label or control signal. To make pose usable by multimodal large language models, we propose a viewpoint-based orientation abstraction that maps orientation into structured viewpoint descriptions while preserving continuous geometric supervision. We further construct an object-centric spatial benchmark (UniSpatial-80K) and train a unified model with an object-token-grounded pose anchor to associate each instance with its pose state. Experiments show that our model improves object-level pose understanding and pose-controllable generation, moving unified models from describing objects toward manipulating spatial states.
Problem

Research questions and friction points this paper is trying to address.

visual understanding
generation
object instances
spatial states
continuous object poses
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Model
Object-Centric
Pose Perception
Viewpoint-based Orientation Abstraction
Controllable Generation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Mining Tan
Mining Tan
Institute of Automation,Chinese Academy of Sciences
Computer VisionMultimediaGenerative AI
Yinuo Wang
Yinuo Wang
Tsinghua University
LLMReinforcement LearningAutonomous DrivingDiffusion Model
Z
Ziqi Zhou
MAIS, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences
Weize Quan
Weize Quan
MAIS-CASIA
Image ProcessingComputer GraphicsDeep Learning
S
Sifei Li
MAIS, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences
J
Jingdong Chen
Ant Group, China
D
DanDan Zheng
Ant Group, China
L
Libin Wang
Ant Group, China
W
Weiming Dong
MAIS, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences