MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对多视角行人检测中未见摄像机配置下的几何视觉捕捉不准确及模型对失真模式依赖的问题,提出了一种基于视觉几何基础模型MV2GF的新方法。
📝 Abstract
Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view map from multi-view images. Recent MVPD methods adopt a unified framework that projects 2D image features into a 3D world space and aggregates them into a single feature. Although they are effective, they struggle to generalize to unseen camera configurations during training due to two main issues. First, they are difficult to capture accurate visual geometry across views in unseen camera configurations. Second, they make detection models highly dependent on distortion patterns during training arising from their image feature projection. To address these, we leverage a visual geometric foundation model and propose MV2GF. This foundation model has exhibited strong generalization in capturing visual geometry across views and predicting accurate 3D attributes in diverse camera configurations. MV2GF fuses task-specific features with general-purpose geometric features extracted by the foundation model to effectively capture the visual geometry even in unseen camera configurations. Furthermore, MV2GF projects each pixel in the image features to an appropriate 3D location using 3D pointmaps predicted by the foundation model, preventing the detection model from depending on distortion patterns during training. Our experiments demonstrate the effectiveness of leveraging a visual geometric foundation model for MVPD and that MV2GF generalizes better than existing methods.
Problem

Research questions and friction points this paper is trying to address.

Multi-View Pedestrian Detection
visual geometry
camera configurations
distortion patterns
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual geometric foundation model
multi-view pedestrian detection
generalization
3D pointmaps
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Taiga Yamane
Human Informatics Laboratories, NTT, Inc., Japan
Satoshi Suzuki
Satoshi Suzuki
NTT
Neural networkscomputer visiondeep learningvideo coding for machines
Ryo Masumura
Ryo Masumura
Distinguished Research Scientist, NTT Corporation
Speech RecognitionSpoken Language ProcessingNatural Language ProcessingComputer Vision
S
Shota Orihashi
Human Informatics Laboratories, NTT, Inc., Japan
T
Tomohiro Tanaka
Human Informatics Laboratories, NTT, Inc., Japan
M
Mana Ihori
Human Informatics Laboratories, NTT, Inc., Japan
N
Naoki Makishima
Human Informatics Laboratories, NTT, Inc., Japan