Developing Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
Current 3D medical imaging AI is hindered by the scarcity of large-scale, paired multimodal datasets, impeding cross-modal alignment and natural language interaction. To address this, we introduce CT-RATE—the first large-scale, paired 3D chest CT–radiology report dataset (25,692 cases)—and propose CT-CLIP, a contrastive learning framework, and CT-CHAT, a vision-language dialogue model. Our contributions include: (1) the first large-scale, fine-grained alignment between 3D CT volumes and free-text radiology reports; (2) CT-CLIP—a task-agnostic foundational model integrating 3D convolutional networks with Vision Transformers, requiring no downstream fine-tuning; and (3) CT-CHAT—the first open-source, 3D CT–specific conversational model, trained via report-driven QA generation and LLM-CT co-fine-tuning for end-to-end diagnostic interaction. Experiments demonstrate that our unsupervised multi-abnormality detection outperforms fully supervised SOTA methods; cross-modal retrieval enables bidirectional image–text queries; and CT-CHAT, fine-tuned on 2.7M medical QA pairs, surpasses existing multimodal medical assistants.