TeaMatch: Teachable Cross-Modal Representation Learning for 2D-3D Matching
This work addresses the challenge of establishing robust 2D–3D correspondences under severe degradation conditions such as noise, low overlap, and structural ambiguity. The authors propose TeaMatch, a novel framework that introduces teachability into cross-modal representation learning for the first time. By employing a task-oriented weak student to simulate typical failure modes and optimizing representations to recover the teacher’s features, TeaMatch enhances structural consistency and robustness. The approach integrates a teacher–student architecture with correspondence-level constraints and geometry-aware regularization, seamlessly fitting into coarse-to-fine matching pipelines without incurring additional inference overhead. Extensive experiments demonstrate that TeaMatch achieves state-of-the-art performance across multiple challenging 2D–3D matching benchmarks, significantly improving matching robustness.