From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种名为Distortion Extenders (DEX)的方法,通过自监督对齐损失最小化来解决广角鱼眼图像的深度估计和开放词汇分割问题,提高了模型在鱼眼图像上的泛化能力。
📝 Abstract
Vision foundation models are capable of generalizing across 3-dimensional (3D) scenes with high-fidelity estimates; their empirical success can be attributed to training on large-scale datasets of perspective images. However, when transferred to wide field-of-view (FoV) images, such as those captured by fisheye cameras, they return erroneous outputs due to a covariate shift stemming from the radial distortion on the image pixels. We propose a method to generalize vision foundation models to fisheye cameras. The crux of our method lies in a set of learnable parameters, termed Distortion Extenders (DEX), that model the fisheye distortion coefficients and the distributional shift between fisheye and perspective images encoded in the latent space. By minimizing a self-supervised alignment loss, DEX transforms the latent embeddings of fisheye images to resemble those of perspective images to recover high-fidelity estimates. DEX is architecture- and task-agnostic: We demonstrate DEX on monocular depth estimation and open-vocabulary segmentation for convolution- and Transformer-based architectures, where we consistently improve over baselines across indoor and outdoor fisheye datasets. As a byproduct, the activations of DEX can also be decoded to distortion coefficients to support camera calibration. Code available at: https://github.com/Suchisrit/DEX.
Problem

Research questions and friction points this paper is trying to address.

fisheye cameras
covariate shift
radial distortion
high-fidelity estimates
vision foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distortion Extenders
self-supervised alignment loss
latent embeddings
monocular depth estimation
open-vocabulary segmentation
💼 Related Jobs
No related jobs found.
R
Rit Gangopadhyay
Yale Vision Laboratory, Yale University, New Haven, CT 06520, USA
Alex Wong
Alex Wong
Yale University
Computer visionMachine learning3D visionUnsupervised learningAdversarial robustness