OPUS: A Simple yet Effective Unified Framework for Open-Vocabulary Detection
本文提出OPUS框架,通过语义丰富的视觉表示和可扩展的定位监督简化开放词汇检测,支持多种提示方式并在多项基准测试中取得领先性能。
本文提出OPUS框架,通过语义丰富的视觉表示和可扩展的定位监督简化开放词汇检测,支持多种提示方式并在多项基准测试中取得领先性能。
This work addresses the high computational cost of high-fidelity 3D generation, which typically relies on large-scale data and models while underutilizing the rich semantic and structural priors embedded in discriminative 3D foundation models. To bridge this gap, the authors propose ROAD, a novel framework that, for the first time, transfers priors from discriminative 3D foundation models into a diffusion Transformer. ROAD introduces a reciprocal objective alignment mechanism that effectively reconciles the heterogeneity between generative and discriminative latent spaces through global semantic compression and optimal micro-structure matching—formulated as bipartite graph matching. Remarkably, without increasing inference overhead, ROAD achieves generation quality on par with the industrial baseline Step1X-3D using only 1.5% of the training data, substantially reducing both training cost and computational requirements.
This work addresses the challenge that existing vision-language models struggle to adapt in in-context learning scenarios where task semantics remain constant but decision criteria dynamically shift. To tackle this, the paper introduces a novel paradigm termed Criterion-Conditioned In-Context Learning (CC-ICL), which requires models to infer implicit decision criteria from provided examples and adjust predictions accordingly. The authors construct the first CC-ICL benchmark, CC-Bench, featuring a dual-layer data structure spanning multiple domains, along with newly proposed metrics—criterion invariance and sensitivity—to evaluate model robustness and adaptability under criterion shifts. Experiments reveal that prevailing models exhibit rigid decision boundary biases, yet simple multi-criterion training substantially enhances criterion sensitivity in 7B-scale models, outperforming closed-source counterparts without compromising general multimodal capabilities.
This work addresses key challenges in large-scale LiDAR scene completion—namely, the loss of geometric detail, spatiotemporal inconsistency, and difficulty in long-range reconstruction—by proposing a diffusion-based generative framework operating on local voxel blocks. The method explicitly models fine-grained geometry within localized 3D regions and introduces two key innovations: a confidence-guided spatiotemporal fusion mechanism and an Annular-Flow diffusion strategy, enabling coherent completion in unbounded space. Evaluated on SemanticKITTI, the model achieves state-of-the-art performance, significantly outperforming existing approaches in both geometric fidelity and temporal consistency. Notably, it demonstrates strong generalization capability by extending a model trained for 20-meter scenes to accurately complete scenes up to 50 meters without any retraining.
Current Driving World Models (DWMs) suffer from limited 3D scene understanding and language-driven reasoning, while point cloud and BEV representations inherently struggle with text-3D alignment. To address these limitations, we propose a 3D Gaussian-based Driving World Model that pioneers early-stage cross-modal alignment by embedding language features into each Gaussian primitive. We introduce a task-aware, language-guided sparse sampling strategy and design a dual-conditioned (language + image) multimodal diffusion architecture for generative modeling. By unifying BEV and point cloud representations, our model enables holistic 3D environmental understanding, language-conditioned scene reasoning, and coherent multimodal generation. Evaluated on nuScenes and NuInteract, it achieves state-of-the-art performance, significantly improving both 3D perception accuracy and generation consistency. The implementation is publicly available.
本文提出OPUS框架,通过语义丰富的视觉表示和可扩展的定位监督简化开放词汇检测,支持多种提示方式并在多项基准测试中取得领先性能。
This work addresses the high computational cost of high-fidelity 3D generation, which typically relies on large-scale data and models while underutilizing the rich semantic and structural priors embedded in discriminative 3D foundation models. To bridge this gap, the authors propose ROAD, a novel framework that, for the first time, transfers priors from discriminative 3D foundation models into a diffusion Transformer. ROAD introduces a reciprocal objective alignment mechanism that effectively reconciles the heterogeneity between generative and discriminative latent spaces through global semantic compression and optimal micro-structure matching—formulated as bipartite graph matching. Remarkably, without increasing inference overhead, ROAD achieves generation quality on par with the industrial baseline Step1X-3D using only 1.5% of the training data, substantially reducing both training cost and computational requirements.
This work addresses the challenge that existing vision-language models struggle to adapt in in-context learning scenarios where task semantics remain constant but decision criteria dynamically shift. To tackle this, the paper introduces a novel paradigm termed Criterion-Conditioned In-Context Learning (CC-ICL), which requires models to infer implicit decision criteria from provided examples and adjust predictions accordingly. The authors construct the first CC-ICL benchmark, CC-Bench, featuring a dual-layer data structure spanning multiple domains, along with newly proposed metrics—criterion invariance and sensitivity—to evaluate model robustness and adaptability under criterion shifts. Experiments reveal that prevailing models exhibit rigid decision boundary biases, yet simple multi-criterion training substantially enhances criterion sensitivity in 7B-scale models, outperforming closed-source counterparts without compromising general multimodal capabilities.
This work addresses key challenges in large-scale LiDAR scene completion—namely, the loss of geometric detail, spatiotemporal inconsistency, and difficulty in long-range reconstruction—by proposing a diffusion-based generative framework operating on local voxel blocks. The method explicitly models fine-grained geometry within localized 3D regions and introduces two key innovations: a confidence-guided spatiotemporal fusion mechanism and an Annular-Flow diffusion strategy, enabling coherent completion in unbounded space. Evaluated on SemanticKITTI, the model achieves state-of-the-art performance, significantly outperforming existing approaches in both geometric fidelity and temporal consistency. Notably, it demonstrates strong generalization capability by extending a model trained for 20-meter scenes to accurately complete scenes up to 50 meters without any retraining.
Current Driving World Models (DWMs) suffer from limited 3D scene understanding and language-driven reasoning, while point cloud and BEV representations inherently struggle with text-3D alignment. To address these limitations, we propose a 3D Gaussian-based Driving World Model that pioneers early-stage cross-modal alignment by embedding language features into each Gaussian primitive. We introduce a task-aware, language-guided sparse sampling strategy and design a dual-conditioned (language + image) multimodal diffusion architecture for generative modeling. By unifying BEV and point cloud representations, our model enables holistic 3D environmental understanding, language-conditioned scene reasoning, and coherent multimodal generation. Evaluated on nuScenes and NuInteract, it achieves state-of-the-art performance, significantly improving both 3D perception accuracy and generation consistency. The implementation is publicly available.