C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection
To address the dual challenges of poor generalization to unseen categories and low robustness under adverse conditions (e.g., low illumination, occlusion) in open-world object detection, this paper proposes a curriculum-based cross-modal contrastive learning framework—first integrating RGB-thermal (RGBT) multimodal perception with vision-language alignment. To mitigate catastrophic forgetting in two-stage training, exponential moving average (EMA) is adopted, providing theoretical guarantees for preserving prior knowledge. Jointly leveraging RGBT pretraining and cross-modal contrastive learning, our method simultaneously enhances category openness and environmental robustness. Extensive experiments on FLIR, OV-COCO, and OV-LVIS benchmarks yield 80.1 AP⁵⁰, 48.6 AP⁵⁰ₙₒᵥₑₗ, and 35.7 mAPᵣ, respectively—outperforming state-of-the-art methods by significant margins.