Multimodal Floorplan Encoding: Learning Dense Modality-Invariant Representations

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态平面图表示学习问题,提出Multimodal Floorplan Encoder (MMFE),通过结合DINOv3和Dense Prediction Transformer及特定训练目标来实现跨模态的几何任务支持。
📝 Abstract
Floorplans arise in many forms, from vector CAD drawings to raster renderings and sensor-derived density maps. This heterogeneity makes it difficult to build learning systems that transfer across modalities and support geometry-centric tasks such as alignment and retrieval. We introduce the Multimodal Floorplan Encoder (MMFE), which maps diverse 2D indoor representations into a shared dense latent grid. MMFE combines a frozen DINOv3 backbone with a trainable Dense Prediction Transformer (DPT) head, and is trained with a per-cell Information Noise-Contrastive Estimation (InfoNCE) objective that aligns spatially corresponding regions across modalities while using all other cells as negatives. To improve robustness to geometric distortions, we incorporate controlled similarity transformations and enforce geometric consistency through feature-grid warping. On Structured3D, a held-out out-of-domain dataset, MMFE improves cross-modal dense matching, enables robust similarity alignment with RANSAC, and yields strong retrieval when paired with learned aggregation.
Problem

Research questions and friction points this paper is trying to address.

Multimodal
Floorplan
Dense Representation
Cross-modal
Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Floorplan Encoder
Dense Prediction Transformer
InfoNCE
🔎 Similar Papers
No similar papers found.
X
Xavier Anadón
University of Zaragoza, Spain
Rémi Pautrat
Rémi Pautrat
Microsoft, ETH Zürich, Ecole polytechnique, INRIA
Computer visiondeep learningrobotics
R
Rui Wang
Microsoft, Switzerland