Institution profile

Niantic

Industry researchnorthamerica · us
Official website
Research library14linked papers
Opportunities0open roles
Selected work

Representative Papers

MVSAnywhere: Zero-Shot Multi-View Stereo

Mar 28, 2025

To address poor generalization, variable input view counts, and unknown effective depth ranges in cross-domain (e.g., indoor-to-outdoor) multi-view depth estimation, this paper proposes a zero-shot cross-scene depth reconstruction method. Our approach introduces an adaptive cost volume fusion mechanism that jointly models monocular priors and multi-view geometric cues; integrates a Transformer-based architecture supporting variable-length view inputs; and employs metadata-driven scale-adaptive cost volume construction and optimization. Crucially, the method requires no target-domain training, accommodates arbitrary numbers of input views, and operates robustly under unknown depth ranges. Evaluated on the Robust Multi-View Depth Benchmark, it achieves state-of-the-art zero-shot performance—significantly outperforming existing monocular and multi-view depth estimation methods—while maintaining architectural flexibility and domain-agnostic inference.

1 citationsRead paper

Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Images

May 19, 2026

This work addresses the challenge of high-quality novel view synthesis in large-scale outdoor scenes, where sparse ground-level imagery leads to insufficient camera coverage. The authors propose a feed-forward method that, for the first time, incorporates orthorectified satellite imagery as a global geometric prior, fusing it with GPS-tagged ground images within a unified geospatial coordinate system. By aligning cross-view features, the method predicts a per-pixel 3D Gaussian splatting representation. This approach substantially improves both scene coverage and rendering quality. Furthermore, the study introduces the first benchmark for georeferenced image-based novel view synthesis, demonstrating superior performance over state-of-the-art methods in terms of reconstruction completeness and visual fidelity.

0 citationsRead paper

NaviNote: Enabling In-situ Spatial Annotation Authoring to Support Exploration and Navigation for Blind and Low Vision People

Mar 09, 2026

This study addresses the limitations of conventional GPS-based geotagging systems, which suffer from insufficient positioning accuracy to support precise navigation for blind and low-vision individuals in unfamiliar environments. To overcome this challenge, the authors propose a novel spatial annotation and navigation system that integrates high-precision visual localization, an intelligent agent architecture, and voice-based interaction. For the first time, centimeter-level visual localization is leveraged to enable accurate last-meter navigation and context-aware spatial annotation for visually impaired users, effectively transcending the precision constraints of traditional GPS. A user study involving 18 blind or low-vision participants demonstrated significant improvements in environmental comprehension and autonomous navigation performance, confirming the system’s effectiveness and innovation as an assistive tool.

0 citationsRead paper

Scene Coordinate Reconstruction Priors

Oct 14, 2025

Insufficient multi-view constraints often degrade scene coordinate regression (SCR) models, adversely affecting downstream 3D vision tasks such as visual relocalization and structure-from-motion (SfM). To address this, we propose a probabilistic training framework that integrates high-order geometric reconstruction priors. Specifically, our method jointly models shallow-depth distributions and leverages a pre-trained 3D point cloud diffusion prior—trained on large-scale indoor scans—to explicitly enforce geometric consistency in predicted scene coordinates. By co-optimizing SCR and point cloud generation during training, the framework significantly improves structural coherence of the learned coordinate space. Evaluated on three indoor benchmarks, our approach achieves more consistent point cloud reconstructions, higher pose estimation success rates, and substantial improvements in novel-view synthesis and camera relocalization performance compared to prior methods.

0 citationsRead paper
Recent publications

Latest Papers

Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Images

May 19, 2026

This work addresses the challenge of high-quality novel view synthesis in large-scale outdoor scenes, where sparse ground-level imagery leads to insufficient camera coverage. The authors propose a feed-forward method that, for the first time, incorporates orthorectified satellite imagery as a global geometric prior, fusing it with GPS-tagged ground images within a unified geospatial coordinate system. By aligning cross-view features, the method predicts a per-pixel 3D Gaussian splatting representation. This approach substantially improves both scene coverage and rendering quality. Furthermore, the study introduces the first benchmark for georeferenced image-based novel view synthesis, demonstrating superior performance over state-of-the-art methods in terms of reconstruction completeness and visual fidelity.

0 citationsRead paper

NaviNote: Enabling In-situ Spatial Annotation Authoring to Support Exploration and Navigation for Blind and Low Vision People

Mar 09, 2026

This study addresses the limitations of conventional GPS-based geotagging systems, which suffer from insufficient positioning accuracy to support precise navigation for blind and low-vision individuals in unfamiliar environments. To overcome this challenge, the authors propose a novel spatial annotation and navigation system that integrates high-precision visual localization, an intelligent agent architecture, and voice-based interaction. For the first time, centimeter-level visual localization is leveraged to enable accurate last-meter navigation and context-aware spatial annotation for visually impaired users, effectively transcending the precision constraints of traditional GPS. A user study involving 18 blind or low-vision participants demonstrated significant improvements in environmental comprehension and autonomous navigation performance, confirming the system’s effectiveness and innovation as an assistive tool.

0 citationsRead paper

Scene Coordinate Reconstruction Priors

Oct 14, 2025

Insufficient multi-view constraints often degrade scene coordinate regression (SCR) models, adversely affecting downstream 3D vision tasks such as visual relocalization and structure-from-motion (SfM). To address this, we propose a probabilistic training framework that integrates high-order geometric reconstruction priors. Specifically, our method jointly models shallow-depth distributions and leverages a pre-trained 3D point cloud diffusion prior—trained on large-scale indoor scans—to explicitly enforce geometric consistency in predicted scene coordinates. By co-optimizing SCR and point cloud generation during training, the framework significantly improves structural coherence of the learned coordinate space. Evaluated on three indoor benchmarks, our approach achieves more consistent point cloud reconstructions, higher pose estimation success rates, and substantial improvements in novel-view synthesis and camera relocalization performance compared to prior methods.

0 citationsRead paper

ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-Training

Oct 13, 2025

Scene Coordinate Regression (SCR) methods suffer from poor generalization in visual relocalization, primarily because conventional frameworks couple training-view information into the regressor’s weights, resulting in low robustness to unseen imaging conditions such as lighting and viewpoint variations. To address this, we propose a novel paradigm that decouples map representation from coordinate regression: we adopt a universal Transformer backbone and inject lightweight, learnable, scene-specific map tokens for each scene. Crucially, we conduct the first large-scale self-supervised pretraining across tens of thousands of scenes—enabling cross-scene knowledge transfer—and subsequently adapt to new scenes via fine-tuning only the map tokens using minimal scene-specific data. Evaluated on multiple challenging relocalization benchmarks, our method significantly improves both robustness and accuracy of pose estimation while maintaining low computational overhead, effectively overcoming the generalization bottleneck of SCR.

0 citationsRead paper