G-ray: Ray-Level Relative Geometric Position Encoding in Multi-View Vision Transformers under Camera Heterogeneity

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究针对多视角视觉Transformer在相机异质性下的相对位置编码问题,提出G-ray方法,通过基于光线角度的旋转相位参数化实现投影不变的位置一致性。
📝 Abstract
We study relative position encoding for multi-view vision Transformers under camera heterogeneity, including varying fields of view (FoVs) or projection models. Existing rotary relative position encodings commonly use image-plane positional coordinates, producing projection-dependent relative phases and inconsistent geometric cues for cross-projection attention. We introduce G-ray, a ray-level relative position encoding whose rotary phases are parameterized by camera-local ray angles. The same camera-local ray pair induces the same relative phase across projections, providing projection-invariant positional consistency. G-ray can be used directly or integrated with existing encodings, retaining complementary geometric cues without additional learned parameters. We validate G-ray in three host encodings, RoPE, GTA, and RayRoPE, across 3D reconstruction and novel-view synthesis (NVS). Across three heterogeneous 3D reconstruction benchmarks at 50 views, G-ray leads all six averaged metrics and reduces mean pointmap relative error by 45.8% over MapAnything, with calibration supplied to both. Trained exclusively on homogeneous pinhole images, the 3D reconstruction model handles mixed pinhole and non-pinhole inputs without retraining and remains competitive on homogeneous pinhole 3D reconstruction protocols. For NVS, GTA and RayRoPE improve with G-ray under joint viewpoint and FoV variation. The project's webpage is available at https://g-ray-project.github.io/.
Problem

Research questions and friction points this paper is trying to address.

relative position encoding
multi-view vision transformers
camera heterogeneity
projection models
geometric cues
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ray-Level Relative Position Encoding
Camera Heterogeneity
Projection-Invariant
Multi-View Vision Transformers
3D Reconstruction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shuo Zhang
Wuhan University, Wuhan, China
X
Xin Su
Wuhan University, Wuhan, China
W
Wei Wang
Wuhan University, Wuhan, China
Jun Liu
Jun Liu
University of Science and Technology of China
X
Xinrui Zeng
Wuhan University, Wuhan, China
Y
Yongsen Chen
Wuhan University, Wuhan, China
C
Chenjie Wang
Rongyun Robot (Guizhou) Co., Ltd.
Guibo Zhu
Guibo Zhu
Institute of Automation, Chinese Academy of Sciecnes
Artificial IntelligenceComputer VisionMachine Learning
J
Jinqiao Wang
Institute of Automation, Chinese Academy of Sciences, Beijing, China, and Wuhan AI Research, Wuhan, China
B
Bin Luo
Wuhan University, Wuhan, China
L
Liangpei Zhang
Wuhan University, Wuhan, China