UniQuery4R: Unified 4D Scene Reconstruction from a Single Query

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
UniQuery4R通过一次编码多帧剪辑并在解码时选择视图,解决了动态4D场景重建中的计算冗余和特征重用问题。
📝 Abstract
Reconstructing dynamic 4D scenes requires jointly estimating correspondence, geometry, object motion, and camera motion. Existing feed-forward methods typically predict dense task-specific maps or independently process source-target pairs, leading to unnecessary computation for sparse queries and limited feature reuse across different frame pairs. We present UniQuery4R, a query-conditioned framework that encodes a multi-frame clip once and selects the source view, target view, and continuous source-image coordinate only at decoding time via source-to-target cross-attention. Each query jointly predicts target correspondence, target-time 3D position, and scene flow, along with source depth, while camera parameters are estimated per view. This design allows the encoded clip to be reused across arbitrary source-target selections and supports both sparse inference and dense reconstruction through batched queries, without learned temporal embeddings tied to a fixed clip length. We further introduce a direction-magnitude parameterization of scene flow with separate supervision for moving and static points. Among the evaluated methods, UniQuery4R achieves the best macro-average results on WorldTrack for both scene-flow estimation and dynamic-point reconstruction.
Problem

Research questions and friction points this paper is trying to address.

4D Scene Reconstruction
Sparse Queries
Feature Reuse
Innovation

Methods, ideas, or system contributions that make the work stand out.

query-conditioned framework
source-to-target cross-attention
direction-magnitude parameterization of scene flow
💼 Related Jobs
No related jobs found.
T
Tiancheng Chen
Kosmo Research
Sheng Tang
Sheng Tang
Institute of Computing Technology, Chinese Academy of Sciences
computer visionpattern recognitionmachine learningimage/video processing
W
Wenhua Jin
Kosmo Research, Automotive Engineering Department, Jilin University
Weiqi Zhang
Weiqi Zhang
Tsinghua University
3D Computer VisionGenerative Model
J
Juntong Fang
School of Software, Tsinghua University
Junsheng Zhou
Junsheng Zhou
Tsinghua University
3D computer vision
Z
Zesong Li
Kosmo Research