Institution profile

MedAI Tech

Industry research
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Decouple and Reason: Anatomically Guided Two-Stage Voxel-Level Grounding of Free-Text Findings in 3D Chest CT

Jul 14, 2026

This work addresses the challenge of precisely aligning free-text descriptions with voxel-level lesions in 3D chest CT scans. To overcome the coupling bottleneck between local feature extraction and semantic understanding inherent in existing end-to-end approaches, the authors propose a decoupled two-stage framework. The first stage performs class-agnostic 3D lesion segmentation, followed by a cross-modal reasoning step that aligns textual descriptions with the segmented regions. The method further incorporates lobar anatomical priors and relative spatial coordinate encoding to enhance spatial disambiguation of local regions. Evaluated on the ReXGroundingCT benchmark, the approach achieves state-of-the-art voxel-level grounding performance, demonstrating the effectiveness of both the decoupled design and the anatomy-guided alignment mechanism.

0 citationsRead paper

Lesion Segmentation in Moderate to Severe Traumatic Brain Injury: An nnU-Net Based Approach with Adaptive Normalization in the AIMS-TBI 2025 Challenge

Jul 14, 2026

This study addresses the significant challenge of segmenting moderate-to-severe traumatic brain injury (msTBI) lesions in T1-weighted MRI, which exhibit high heterogeneity in size, shape, and location. The authors propose an anatomically constrained adaptive intensity normalization method that operates exclusively within the brain parenchyma, effectively reducing inter-subject intensity variability and suppressing non-brain tissue artifacts. Integrated with the nnU-Net deep learning framework, this strategy substantially enhances lesion segmentation robustness. Evaluated on the AIMS-TBI 2025 Challenge test set, the model achieved a global Dice score of 0.6305, a lesion Dice of 0.4805, and a notably high non-lesion Dice of 0.9324, demonstrating strong specificity and promising clinical applicability.

0 citationsRead paper

SAM 3D for 3D Object Reconstruction from Remote Sensing Images

Dec 26, 2025

Monocular remote sensing image-based 3D building reconstruction is hindered by reliance on task-specific architectures and dense, labor-intensive supervision. Method: This paper presents the first systematic evaluation and adaptation of the general-purpose image-to-3D foundation model SAM 3D to remote sensing. We propose a “segment–reconstruct–compose” pipeline for structured, city-scale 3D modeling and introduce the first SAM extension tailored for remote sensing 3D reconstruction. To objectively assess geometric fidelity, we propose CLIP-based Multi-Modal Distance (CMMD), a novel metric quantifying reconstruction quality. Contribution/Results: Evaluated on the NYC Urban Dataset, our approach significantly outperforms TRELLIS, yielding more coherent roof geometries with sharper boundaries. Both FID and CMMD scores show substantial improvement. This work breaks the dependency on task-specific designs and empirically validates the feasibility and effectiveness of leveraging general-purpose foundation models for large-scale urban 3D scene reconstruction.

0 citationsRead paper
Recent publications

Latest Papers

Decouple and Reason: Anatomically Guided Two-Stage Voxel-Level Grounding of Free-Text Findings in 3D Chest CT

Jul 14, 2026

This work addresses the challenge of precisely aligning free-text descriptions with voxel-level lesions in 3D chest CT scans. To overcome the coupling bottleneck between local feature extraction and semantic understanding inherent in existing end-to-end approaches, the authors propose a decoupled two-stage framework. The first stage performs class-agnostic 3D lesion segmentation, followed by a cross-modal reasoning step that aligns textual descriptions with the segmented regions. The method further incorporates lobar anatomical priors and relative spatial coordinate encoding to enhance spatial disambiguation of local regions. Evaluated on the ReXGroundingCT benchmark, the approach achieves state-of-the-art voxel-level grounding performance, demonstrating the effectiveness of both the decoupled design and the anatomy-guided alignment mechanism.

0 citationsRead paper

Lesion Segmentation in Moderate to Severe Traumatic Brain Injury: An nnU-Net Based Approach with Adaptive Normalization in the AIMS-TBI 2025 Challenge

Jul 14, 2026

This study addresses the significant challenge of segmenting moderate-to-severe traumatic brain injury (msTBI) lesions in T1-weighted MRI, which exhibit high heterogeneity in size, shape, and location. The authors propose an anatomically constrained adaptive intensity normalization method that operates exclusively within the brain parenchyma, effectively reducing inter-subject intensity variability and suppressing non-brain tissue artifacts. Integrated with the nnU-Net deep learning framework, this strategy substantially enhances lesion segmentation robustness. Evaluated on the AIMS-TBI 2025 Challenge test set, the model achieved a global Dice score of 0.6305, a lesion Dice of 0.4805, and a notably high non-lesion Dice of 0.9324, demonstrating strong specificity and promising clinical applicability.

0 citationsRead paper

SAM 3D for 3D Object Reconstruction from Remote Sensing Images

Dec 26, 2025

Monocular remote sensing image-based 3D building reconstruction is hindered by reliance on task-specific architectures and dense, labor-intensive supervision. Method: This paper presents the first systematic evaluation and adaptation of the general-purpose image-to-3D foundation model SAM 3D to remote sensing. We propose a “segment–reconstruct–compose” pipeline for structured, city-scale 3D modeling and introduce the first SAM extension tailored for remote sensing 3D reconstruction. To objectively assess geometric fidelity, we propose CLIP-based Multi-Modal Distance (CMMD), a novel metric quantifying reconstruction quality. Contribution/Results: Evaluated on the NYC Urban Dataset, our approach significantly outperforms TRELLIS, yielding more coherent roof geometries with sharper boundaries. Both FID and CMMD scores show substantial improvement. This work breaks the dependency on task-specific designs and empirically validates the feasibility and effectiveness of leveraging general-purpose foundation models for large-scale urban 3D scene reconstruction.

0 citationsRead paper