Institution profile

COWAROBOT

Industry researchasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models

Mar 10, 2026

This work addresses the challenge of accurate localization from natural language descriptions to 3D point cloud maps by proposing VLM-Loc, a novel approach that leverages vision-language models (VLMs) for cross-modal alignment between text and point clouds. By transforming point clouds into bird’s-eye-view images and structured scene graphs, the method jointly encodes geometric and semantic information. A partial node matching mechanism is introduced to enable interpretable spatial reasoning. Evaluated on the newly constructed CityLoc benchmark, VLM-Loc significantly outperforms existing methods, achieving state-of-the-art performance in both localization accuracy and robustness, while simultaneously enhancing the model’s spatial reasoning capabilities and decision interpretability.

0 citationsRead paper

Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs

Aug 30, 2025

Multimodal large language models (MLLMs) suffer from object hallucination, comprising two distinct types: omission hallucination (failing to describe present objects) and fabrication hallucination (describing non-existent objects). Prior methods erroneously assume a shared origin, leading to trade-offs between the two. This work is the first to identify their heterogeneous causes: omission stems from insufficient confidence in visual-to-linguistic mapping, whereas fabrication arises from spurious cross-modal associations in the joint representation space. To address this, we propose the Visual-Semantic Attention Potential Field—a theoretical framework—and design VPFC, a plug-and-play, fine-tuning-free calibration method. VPFC achieves decoupled hallucination control via visual attention intervention, statistical bias analysis, and cross-modal association disentanglement. Experiments demonstrate that VPFC simultaneously reduces omission hallucination and suppresses fabrication hallucination, establishing a novel paradigm for MLLM hallucination mitigation—interpretable, balanced, and robust.

0 citationsRead paper

You Only Click Once: Single Point Weakly Supervised 3D Instance Segmentation for Autonomous Driving

Feb 27, 2025

To address the high annotation cost in LiDAR point cloud 3D instance segmentation for autonomous driving, this paper proposes a single-click weakly supervised paradigm: only one click on the bird’s-eye view (BEV) is required to generate high-quality 3D pseudo-labels. Methodologically, we introduce the first weakly supervised framework integrating vision foundation models with point cloud geometric constraints, incorporating cross-frame temporal consistency modeling, density-aware spatial modeling, and IoU-confidence-guided collaborative pseudo-label refinement. On the Waymo Open Dataset, our method achieves performance comparable to fully supervised Cylinder3D using merely 0.8% of full annotations—significantly outperforming existing weakly supervised approaches and establishing new state-of-the-art (SOTA) results. Our core contributions are threefold: (1) the first formalization of the single-click weak supervision setting for 3D instance segmentation; (2) a novel geometry–semantics joint modeling paradigm; and (3) an efficient, robust pseudo-label generation and optimization mechanism.

0 citationsRead paper

Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving

Jan 15, 2025

To address three core challenges in end-to-end autonomous driving—weak visual understanding, difficult decision reasoning, and poor scene generalization—this paper proposes GPVL, a unified vision-language framework. GPVL introduces the first 3D bird’s-eye-view (BEV) vision-language joint pretraining paradigm: multi-view images are encoded into BEV features and explicitly aligned with language representations. A cross-modal language model is then designed to autoregressively generate both high-level driving instructions and low-level fine-grained trajectories. By unifying perception, comprehension, and planning, GPVL enables full-stack autonomous driving conditioned on natural language commands. Evaluated on nuScenes, GPVL significantly outperforms state-of-the-art methods across key metrics, demonstrates strong generalization to unseen scenarios, and exhibits promising potential for real-time deployment.

0 citationsRead paper
Recent publications

Latest Papers

VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models

Mar 10, 2026

This work addresses the challenge of accurate localization from natural language descriptions to 3D point cloud maps by proposing VLM-Loc, a novel approach that leverages vision-language models (VLMs) for cross-modal alignment between text and point clouds. By transforming point clouds into bird’s-eye-view images and structured scene graphs, the method jointly encodes geometric and semantic information. A partial node matching mechanism is introduced to enable interpretable spatial reasoning. Evaluated on the newly constructed CityLoc benchmark, VLM-Loc significantly outperforms existing methods, achieving state-of-the-art performance in both localization accuracy and robustness, while simultaneously enhancing the model’s spatial reasoning capabilities and decision interpretability.

0 citationsRead paper

Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs

Aug 30, 2025

Multimodal large language models (MLLMs) suffer from object hallucination, comprising two distinct types: omission hallucination (failing to describe present objects) and fabrication hallucination (describing non-existent objects). Prior methods erroneously assume a shared origin, leading to trade-offs between the two. This work is the first to identify their heterogeneous causes: omission stems from insufficient confidence in visual-to-linguistic mapping, whereas fabrication arises from spurious cross-modal associations in the joint representation space. To address this, we propose the Visual-Semantic Attention Potential Field—a theoretical framework—and design VPFC, a plug-and-play, fine-tuning-free calibration method. VPFC achieves decoupled hallucination control via visual attention intervention, statistical bias analysis, and cross-modal association disentanglement. Experiments demonstrate that VPFC simultaneously reduces omission hallucination and suppresses fabrication hallucination, establishing a novel paradigm for MLLM hallucination mitigation—interpretable, balanced, and robust.

0 citationsRead paper

You Only Click Once: Single Point Weakly Supervised 3D Instance Segmentation for Autonomous Driving

Feb 27, 2025

To address the high annotation cost in LiDAR point cloud 3D instance segmentation for autonomous driving, this paper proposes a single-click weakly supervised paradigm: only one click on the bird’s-eye view (BEV) is required to generate high-quality 3D pseudo-labels. Methodologically, we introduce the first weakly supervised framework integrating vision foundation models with point cloud geometric constraints, incorporating cross-frame temporal consistency modeling, density-aware spatial modeling, and IoU-confidence-guided collaborative pseudo-label refinement. On the Waymo Open Dataset, our method achieves performance comparable to fully supervised Cylinder3D using merely 0.8% of full annotations—significantly outperforming existing weakly supervised approaches and establishing new state-of-the-art (SOTA) results. Our core contributions are threefold: (1) the first formalization of the single-click weak supervision setting for 3D instance segmentation; (2) a novel geometry–semantics joint modeling paradigm; and (3) an efficient, robust pseudo-label generation and optimization mechanism.

0 citationsRead paper

Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving

Jan 15, 2025

To address three core challenges in end-to-end autonomous driving—weak visual understanding, difficult decision reasoning, and poor scene generalization—this paper proposes GPVL, a unified vision-language framework. GPVL introduces the first 3D bird’s-eye-view (BEV) vision-language joint pretraining paradigm: multi-view images are encoded into BEV features and explicitly aligned with language representations. A cross-modal language model is then designed to autoregressively generate both high-level driving instructions and low-level fine-grained trajectories. By unifying perception, comprehension, and planning, GPVL enables full-stack autonomous driving conditioned on natural language commands. Evaluated on nuScenes, GPVL significantly outperforms state-of-the-art methods across key metrics, demonstrates strong generalization to unseen scenarios, and exhibits promising potential for real-time deployment.

0 citationsRead paper