Tele360: Real-Time Feed-Forward Human Reconstruction from Sparse Unposed Cameras

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Tele360通过轻量级多视图变压器和高斯解码器,实现实时从稀疏未定位RGB流中重建动态人物,并支持2K分辨率下的实时渲染。
📝 Abstract
Live free-viewpoint visualization of real humans is critical for immersive communication and interactive digital experiences. Existing methods either rely on computationally expensive optimization or require calibrated cameras and low-resolution inputs, making real-time high-resolution deployment impractical. In this work, we present Tele360, the first real-time feed-forward system for dynamic human reconstruction and live free-viewpoint visualization from sparse, unposed RGB streams. Our system jointly estimates camera poses and reconstructs a dynamic 3D Gaussian representation for each time instance in a single forward pass. To achieve this, we start by designing a lightweight sparsity-aware multi-view transformer backbone that tokenizes foreground human regions while preserving global context through a shared scene token. We then employ a fully transformer-based Gaussian decoder to mitigate convolution-induced over-smoothing while keeping decoding sparse and efficient. In addition, we introduce a hybrid feature pyramid that injects multi-scale appearance cues into geometry prediction. We further introduce a lightweight differentiable Levenberg-Marquardt camera refinement layer to enhance multi-view consistency and geometric alignment. Moreover, to stabilize learning under sparse, unposed inputs, we transfer multi-view geometry priors from a large visual-geometry foundation model via teacher-student distillation. Finally, the predicted Gaussian maps are streamed with video codecs to remote devices for interactive free-viewpoint rendering. Extensive experiments show that Tele360 achieves state-of-the-art visual quality on studio benchmarks while supporting real-time 2K input-to-rendering at over 25 FPS on a single consumer GPU. Additional captured sequences illustrate its performance across varied subjects, clothing, and motions under our multi-camera setup.
Problem

Research questions and friction points this paper is trying to address.

real-time
free-viewpoint visualization
dynamic human reconstruction
sparse unposed cameras
high-resolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

real-time feed-forward system
dynamic 3D Gaussian representation
sparsity-aware multi-view transformer
transformer-based Gaussian decoder
camera refinement layer
🔎 Similar Papers
No similar papers found.