Institution profile

Insta360

Industry researchasia · cn
Official website
Research library12linked papers
Opportunities0open roles
Selected work

Representative Papers

PanoWorld: Real-World Panoramic Generation

Jul 10, 2026

This work addresses the challenge of modeling long-range memory while preserving physical consistency in panoramic world models under large-scale spatial variations and complex lighting conditions. The authors propose a novel approach based on rotation-equivariant representations that simplifies camera trajectories to translational motion under a fixed orientation. By jointly modeling current actions and historical memory through Dense Panoramic Ray Conditioning (DPRC) and a Geometry-aware Memory Augmentation (GMA) mechanism, the method leverages a three-stage progressive training strategy to optimize the entire system. Key contributions include the first incorporation of rotation equivariance into panoramic world modeling to reduce trajectory complexity, the introduction of World360—the first large-scale hybrid real-simulated panoramic dataset—and state-of-the-art performance on this benchmark, demonstrating superior physical consistency and generation quality.

0 citationsRead paper

Geometry and Gradient-based Partitioning for Panoramic Outdoor Reconstruction

Jul 09, 2026

This work addresses the challenges of high data and computational costs, as well as the failure of conventional frustum-based partitioning strategies, in 3D Gaussian Splatting (3DGS) reconstruction of large-scale outdoor panoramic scenes. The authors propose PanoLOG, a coarse-to-fine two-stage adaptive tiling framework that introduces a novel geometry- and gradient-driven partitioning strategy (G²PS). This approach integrates celestial sphere modeling, monocular depth supervision, disparity-driven uncertainty estimation, and gradient-based importance scoring to enable efficient voxelization and camera assignment. The study also presents Pano360, the first large-scale outdoor panoramic 3DGS benchmark dataset, achieving state-of-the-art rendering quality while maintaining parallel scalability. All code, models, and datasets are publicly released.

0 citationsRead paper

Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control

May 28, 2026

This work addresses the challenge that standard sampling and test-time guidance in diffusion models for visual navigation often produce trajectory updates that deviate from the training data manifold, yielding unreliable or inefficient paths. The authors propose a training-free inference method that combines Fisher information-preserving guidance with subspace projection via outer products to optimize task objectives while suppressing out-of-distribution actions that induce Fisher drift. Introducing, for the first time, Fisher-preserving updates and truncated Fisher denoising sensitivity as uncertainty-aware metrics, the approach enables robust fusion of multi-sample actions. Computational efficiency is further enhanced through low-rank Jacobian decomposition and single-step backpropagation. Evaluated on Maze2D, PushT, and real-world robotic visual navigation benchmarks, the method significantly outperforms existing diffusion-based policies, markedly improving trajectory reliability and task success rates.

0 citationsRead paper

MOSIV: Multi-Object System Identification from Videos

Mar 06, 2026

This work addresses the challenge of estimating continuous material parameters from videos involving multiple interacting objects—a setting where existing methods, often limited to single-object scenarios or discrete material classifications, struggle to generalize. To this end, we propose MOSIV, a novel framework that, for the first time, enables joint estimation of per-object continuous material properties from multi-object interaction videos rich in contact dynamics. MOSIV integrates differentiable physics simulation with a geometry-aware alignment objective and introduces object-level fine-grained supervision, substantially improving optimization stability and estimation accuracy. Evaluated on a newly constructed synthetic benchmark featuring complex multi-object interactions, MOSIV significantly outperforms current approaches in both system identification accuracy and long-horizon simulation fidelity.

0 citationsRead paper

Seg-VAR: Image Segmentation with Visual Autoregressive Modeling

Nov 16, 2025

Traditional discriminative models for image segmentation suffer from insufficient modeling of low-level spatial structures. To address this, we propose Seg-VAR—the first generative segmentation framework that introduces Vision Autoregressive (VAR) modeling into dense prediction tasks. Our method reformulates segmentation as a conditional mask generation problem: an image encoder and a spatially aware seglat encoder jointly align the image–mask latent distributions, while a position-sensitive color mapping discretizes masks into sequence tokens to enable hierarchical, fine-grained prediction. A multi-stage training strategy enhances stability in implicit representation learning. Evaluated on PASCAL VOC and COCO-Stuff benchmarks, Seg-VAR significantly outperforms state-of-the-art discriminative and generative approaches. These results demonstrate the effectiveness and scalability of autoregressive modeling for pixel-level generation tasks.

0 citationsRead paper
Recent publications

Latest Papers

PanoWorld: Real-World Panoramic Generation

Jul 10, 2026

This work addresses the challenge of modeling long-range memory while preserving physical consistency in panoramic world models under large-scale spatial variations and complex lighting conditions. The authors propose a novel approach based on rotation-equivariant representations that simplifies camera trajectories to translational motion under a fixed orientation. By jointly modeling current actions and historical memory through Dense Panoramic Ray Conditioning (DPRC) and a Geometry-aware Memory Augmentation (GMA) mechanism, the method leverages a three-stage progressive training strategy to optimize the entire system. Key contributions include the first incorporation of rotation equivariance into panoramic world modeling to reduce trajectory complexity, the introduction of World360—the first large-scale hybrid real-simulated panoramic dataset—and state-of-the-art performance on this benchmark, demonstrating superior physical consistency and generation quality.

0 citationsRead paper

Geometry and Gradient-based Partitioning for Panoramic Outdoor Reconstruction

Jul 09, 2026

This work addresses the challenges of high data and computational costs, as well as the failure of conventional frustum-based partitioning strategies, in 3D Gaussian Splatting (3DGS) reconstruction of large-scale outdoor panoramic scenes. The authors propose PanoLOG, a coarse-to-fine two-stage adaptive tiling framework that introduces a novel geometry- and gradient-driven partitioning strategy (G²PS). This approach integrates celestial sphere modeling, monocular depth supervision, disparity-driven uncertainty estimation, and gradient-based importance scoring to enable efficient voxelization and camera assignment. The study also presents Pano360, the first large-scale outdoor panoramic 3DGS benchmark dataset, achieving state-of-the-art rendering quality while maintaining parallel scalability. All code, models, and datasets are publicly released.

0 citationsRead paper

Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control

May 28, 2026

This work addresses the challenge that standard sampling and test-time guidance in diffusion models for visual navigation often produce trajectory updates that deviate from the training data manifold, yielding unreliable or inefficient paths. The authors propose a training-free inference method that combines Fisher information-preserving guidance with subspace projection via outer products to optimize task objectives while suppressing out-of-distribution actions that induce Fisher drift. Introducing, for the first time, Fisher-preserving updates and truncated Fisher denoising sensitivity as uncertainty-aware metrics, the approach enables robust fusion of multi-sample actions. Computational efficiency is further enhanced through low-rank Jacobian decomposition and single-step backpropagation. Evaluated on Maze2D, PushT, and real-world robotic visual navigation benchmarks, the method significantly outperforms existing diffusion-based policies, markedly improving trajectory reliability and task success rates.

0 citationsRead paper

MOSIV: Multi-Object System Identification from Videos

Mar 06, 2026

This work addresses the challenge of estimating continuous material parameters from videos involving multiple interacting objects—a setting where existing methods, often limited to single-object scenarios or discrete material classifications, struggle to generalize. To this end, we propose MOSIV, a novel framework that, for the first time, enables joint estimation of per-object continuous material properties from multi-object interaction videos rich in contact dynamics. MOSIV integrates differentiable physics simulation with a geometry-aware alignment objective and introduces object-level fine-grained supervision, substantially improving optimization stability and estimation accuracy. Evaluated on a newly constructed synthetic benchmark featuring complex multi-object interactions, MOSIV significantly outperforms current approaches in both system identification accuracy and long-horizon simulation fidelity.

0 citationsRead paper

Seg-VAR: Image Segmentation with Visual Autoregressive Modeling

Nov 16, 2025

Traditional discriminative models for image segmentation suffer from insufficient modeling of low-level spatial structures. To address this, we propose Seg-VAR—the first generative segmentation framework that introduces Vision Autoregressive (VAR) modeling into dense prediction tasks. Our method reformulates segmentation as a conditional mask generation problem: an image encoder and a spatially aware seglat encoder jointly align the image–mask latent distributions, while a position-sensitive color mapping discretizes masks into sequence tokens to enable hierarchical, fine-grained prediction. A multi-stage training strategy enhances stability in implicit representation learning. Evaluated on PASCAL VOC and COCO-Stuff benchmarks, Seg-VAR significantly outperforms state-of-the-art discriminative and generative approaches. These results demonstrate the effectiveness and scalability of autoregressive modeling for pixel-level generation tasks.

0 citationsRead paper