Advancing AI-Powered Medical Image Synthesis: Insights from MedVQA-GI Challenge Using CLIP, Fine-Tuned Stable Diffusion, and Dream-Booth + LoRA
This work addresses the critical bottleneck in medical diagnosis: the absence of methods for dynamically generating high-fidelity images from clinical text. We propose the first dual-task text-to-image framework tailored for gastrointestinal imaging—comprising Image Synthesis (IS) and Optimal Prompt Generation (OPG). Methodologically, we systematically integrate fine-tuned Stable Diffusion, DreamBooth-based personalization, and LoRA-based low-rank adaptation, coupled with a CLIP text encoder to jointly optimize generation quality, class controllability, and diversity. Evaluated on multi-center data, our approach achieves FID = 0.064 and Inception Score = 2.327, significantly outperforming baseline models. Key contributions include: (1) transcending traditional static image analysis by enabling dynamic, natural-language-driven medical image synthesis; (2) establishing a scalable, high-precision prompt optimization mechanism; and (3) providing a reproducible technical pipeline and standardized evaluation benchmark for clinical-oriented generative AI.