View Selection for 3D Captioning via Diffusion Ranking
To address vision-language hallucinations in 3D object captioning caused by rendering viewpoint mismatch, this paper proposes an unsupervised view-ranking method based on diffusion models. Specifically, we leverage pre-trained text-to-3D models (e.g., Stable Diffusion 3D) to quantify 3D–2D view alignment scores and select the most discriminative 2D views for input to multimodal large language models (e.g., GPT-4V) to generate accurate captions. This approach extends the view-ranking paradigm to 3D visual question answering (3D-VQA), achieving significant improvements over CLIP-based baselines on Objaverse/XL. Furthermore, we correct 200K erroneous captions in Cap3D and construct the first million-scale, high-quality 3D-caption dataset—Cap3D-v2—establishing a robust benchmark for 3D understanding and generation.