RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching
为解决AI生成的放射学报告临床质量评估难题,提出RadMatch方法,通过结构化发现匹配和多维度评分来提供可解释和可审计的评估结果。
为解决AI生成的放射学报告临床质量评估难题,提出RadMatch方法,通过结构化发现匹配和多维度评分来提供可解释和可审计的评估结果。
This work addresses the limited generalization of existing medical foundation models in 3D dense prediction tasks, where frozen encoders significantly underperform fully trained nnU-Net. To overcome this, the authors propose an enhanced convolutional Masked Autoencoder (MAE) pretraining framework incorporating robust reconstruction targets, feature regularization, and local-global similarity-based contrastive learning. For the first time, this approach enables unified pretraining across large-scale, multimodal (CT/MRI) datasets spanning diverse anatomical regions. Evaluated on eight segmentation benchmarks, the method substantially outperforms strong MAE baselines: with a frozen encoder, it achieves markedly improved performance, while remaining competitive under full fine-tuning—particularly excelling in lesion segmentation tasks with scarce annotations. These results demonstrate its potential to support efficient clinical deployment.
Existing SAM-based methods operate on 2D images and cannot directly process 3D medical volumes (e.g., CT/MRI); moreover, autoregressive slice-wise inference and sliding-window strategies incur high computational overhead, latency, and implementation complexity. To address these limitations, we propose an end-to-end, lightweight 3D prompt-driven segmentation framework built upon the SegVol architecture. Our method introduces a compact 3D prompt encoder and a voxel-wise fully convolutional decoder, enabling direct support for interactive prompts—including points and bounding boxes—without sliding windows or sequential modeling. Evaluated across multiple 3D medical imaging benchmarks, our approach achieves state-of-the-art segmentation accuracy while accelerating inference by 2.1–3.8× and reducing GPU memory consumption by 47%–63%. These improvements significantly enhance real-time performance and clinical usability for interactive 3D medical image segmentation.
To address the limitations of slice-wise prompting and lack of interactive editing in 3D medical image segmentation, this paper proposes the first single-prompt-driven framework for 3D segmentation using a 2D foundation model. Methodologically, it introduces (1) noisy masks as a novel weakly supervised prompt type; (2) a slice-wise iterative inference mechanism jointly optimized with 3D consistency constraints; and (3) the first benchmark supporting cross-domain generalization and real-time interactive editing evaluation for single-prompt 3D medical segmentation. Built upon the SAM architecture, the framework fuses multimodal prompts—including noisy masks, points, and bounding boxes—to achieve high-precision 3D organ segmentation from a single prompt on the AMOS dataset, attaining an mDice of 82.7%, substantially outperforming existing methods. It further demonstrates strong cross-domain transferability and clinically viable interactive editing capabilities.
为解决AI生成的放射学报告临床质量评估难题,提出RadMatch方法,通过结构化发现匹配和多维度评分来提供可解释和可审计的评估结果。
This work addresses the limited generalization of existing medical foundation models in 3D dense prediction tasks, where frozen encoders significantly underperform fully trained nnU-Net. To overcome this, the authors propose an enhanced convolutional Masked Autoencoder (MAE) pretraining framework incorporating robust reconstruction targets, feature regularization, and local-global similarity-based contrastive learning. For the first time, this approach enables unified pretraining across large-scale, multimodal (CT/MRI) datasets spanning diverse anatomical regions. Evaluated on eight segmentation benchmarks, the method substantially outperforms strong MAE baselines: with a frozen encoder, it achieves markedly improved performance, while remaining competitive under full fine-tuning—particularly excelling in lesion segmentation tasks with scarce annotations. These results demonstrate its potential to support efficient clinical deployment.
Existing SAM-based methods operate on 2D images and cannot directly process 3D medical volumes (e.g., CT/MRI); moreover, autoregressive slice-wise inference and sliding-window strategies incur high computational overhead, latency, and implementation complexity. To address these limitations, we propose an end-to-end, lightweight 3D prompt-driven segmentation framework built upon the SegVol architecture. Our method introduces a compact 3D prompt encoder and a voxel-wise fully convolutional decoder, enabling direct support for interactive prompts—including points and bounding boxes—without sliding windows or sequential modeling. Evaluated across multiple 3D medical imaging benchmarks, our approach achieves state-of-the-art segmentation accuracy while accelerating inference by 2.1–3.8× and reducing GPU memory consumption by 47%–63%. These improvements significantly enhance real-time performance and clinical usability for interactive 3D medical image segmentation.
To address the limitations of slice-wise prompting and lack of interactive editing in 3D medical image segmentation, this paper proposes the first single-prompt-driven framework for 3D segmentation using a 2D foundation model. Methodologically, it introduces (1) noisy masks as a novel weakly supervised prompt type; (2) a slice-wise iterative inference mechanism jointly optimized with 3D consistency constraints; and (3) the first benchmark supporting cross-domain generalization and real-time interactive editing evaluation for single-prompt 3D medical segmentation. Built upon the SAM architecture, the framework fuses multimodal prompts—including noisy masks, points, and bounding boxes—to achieve high-precision 3D organ segmentation from a single prompt on the AMOS dataset, attaining an mDice of 82.7%, substantially outperforming existing methods. It further demonstrates strong cross-domain transferability and clinically viable interactive editing capabilities.