CNsEMD: An Expert-Annotated Multi-Field-Strength MRI Dataset and a Hyperspherical Manifold Network for Multimodal Cranial Nerve Parcellation
为解决颅神经分割难题,本文引入了专家标注的多模态数据集CNsEMD,并提出了一种超球面流形网络PHM-Net来学习跨模态表示。
为解决颅神经分割难题,本文引入了专家标注的多模态数据集CNsEMD,并提出了一种超球面流形网络PHM-Net来学习跨模态表示。
This study addresses the challenges of scarce expert annotations and modality heterogeneity in diagnosing plaque vulnerability from 3D carotid MRI. We propose a text-guided variational multimodal knowledge distillation (VMD) framework. Methodologically, it integrates a 3D CNN with a text encoder and employs variational inference to explicitly model predictive uncertainty, while leveraging domain knowledge embedded in radiology reports to enable cross-modal image–text alignment and knowledge transfer. Our key contribution lies in embedding expert priors into the variational learning paradigm, enabling robust diagnosis of unlabeled imaging data under minimal supervision. Evaluated on a proprietary clinical dataset, VMD significantly outperforms unimodal and conventional multimodal baselines: with only 5% labeled data, it achieves 92.3% diagnostic accuracy. The framework establishes a novel, interpretable, and generalizable multimodal learning paradigm for low-resource medical image analysis.
为解决颅神经分割难题,本文引入了专家标注的多模态数据集CNsEMD,并提出了一种超球面流形网络PHM-Net来学习跨模态表示。
This study addresses the challenges of scarce expert annotations and modality heterogeneity in diagnosing plaque vulnerability from 3D carotid MRI. We propose a text-guided variational multimodal knowledge distillation (VMD) framework. Methodologically, it integrates a 3D CNN with a text encoder and employs variational inference to explicitly model predictive uncertainty, while leveraging domain knowledge embedded in radiology reports to enable cross-modal image–text alignment and knowledge transfer. Our key contribution lies in embedding expert priors into the variational learning paradigm, enabling robust diagnosis of unlabeled imaging data under minimal supervision. Evaluated on a proprietary clinical dataset, VMD significantly outperforms unimodal and conventional multimodal baselines: with only 5% labeled data, it achieves 92.3% diagnostic accuracy. The framework establishes a novel, interpretable, and generalizable multimodal learning paradigm for low-resource medical image analysis.