Decouple and Reason: Anatomically Guided Two-Stage Voxel-Level Grounding of Free-Text Findings in 3D Chest CT
This work addresses the challenge of precisely aligning free-text descriptions with voxel-level lesions in 3D chest CT scans. To overcome the coupling bottleneck between local feature extraction and semantic understanding inherent in existing end-to-end approaches, the authors propose a decoupled two-stage framework. The first stage performs class-agnostic 3D lesion segmentation, followed by a cross-modal reasoning step that aligns textual descriptions with the segmented regions. The method further incorporates lobar anatomical priors and relative spatial coordinate encoding to enhance spatial disambiguation of local regions. Evaluated on the ReXGroundingCT benchmark, the approach achieves state-of-the-art voxel-level grounding performance, demonstrating the effectiveness of both the decoupled design and the anatomy-guided alignment mechanism.