FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the time-consuming nature of pulmonary nodule assessment and the lack of unified clinical interpretability frameworks in existing AI systems by proposing the first two-stage vision-language model for this domain. By fine-tuning Florence-2 for imaging attribute extraction and integrating Zephyr-7B to generate structured descriptions with follow-up recommendations, the framework achieves end-to-end clinical decision support. Experimental results demonstrate that the proposed method outperforms both GPT-4V and human baselines in attribute recognition, achieving an expert-evaluated comprehensive score of 89.5% while ensuring clinical safety. Consequently, this work effectively bridges the critical gap in interpretable intelligent analysis for pulmonary nodules, offering a robust solution that aligns automated assessments with standardized clinical reasoning protocols.
📝 Abstract
Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment. After pulmonary nodule detection, radiologists manually assess anatomical location, diameter, margin characteristics, and attenuation type to support risk assessment and clinical decision-making. However, this post-detection workflow is time-consuming and can be affected by inter-observer variability. Existing Artificial Intelligence methods often focus on isolated tasks, limiting their use as a unified, clinically grounded interpretation framework. This study presents FZ-VLM, a two-stage Florence-Zephyr Vision Language Model framework for unified structured pulmonary nodule characterization in lung CT. The framework uses a fine-tuned Florence-2 model to extract radiological attributes from expert-annotated 2D axial CT slices, while a Zephyr-7B model uses these attributes to generate nodule descriptions, follow-up recommendations, and longitudinal analyses. Results showed that the Stage 1 model achieved 77.18\% accuracy for anatomical location, 67.96\% accuracy for margin characteristics, and 79.13\% accuracy for attenuation type, with a Mean Absolute Error of 2.58 mm for diameter estimation, outperforming evaluated GPT-4-based baselines as well as the human baseline. Expert radiologist evaluation of Stage 2 showed 93.9\% accuracy, 98.6\% completeness score, 76.1\% clinical relevance, and an overall score of 89.5\%. Safety analysis showed that most outputs were clinically safe, although some follow-up recommendations still required expert review. To the best of our knowledge, this study presents the first two-stage Vision-Language Model framework for structured nodule characterization and clinical decision-making.
Problem

Research questions and friction points this paper is trying to address.

Pulmonary Nodule Characterization
Clinical Decision Making
Vision Language Model
Lung CT
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Model
Two-stage Framework
Pulmonary Nodule Characterization
Clinical Decision Making
Florence-Zephyr
💼 Related Jobs
No related jobs found.
P
Pramit Dutta
College of Engineering, University of Guelph, Guelph, ON N1G 2W1, Canada
J
Jenita Manokaran
Holland Bloorview Kids Rehabilitation Hospital, Toronto, ON M4G 1R8, Canada
R
Richa Mittal
Guelph General Hospital, Guelph, ON N1E 4J4, Canada
R
Ryan Appleby
Department of Clinical Studies, Ontario Veterinary College, University of Guelph, Guelph, ON N1G 2W1, Canada
Eranga Ukwatta
Eranga Ukwatta
Associate Professor, School of Engineering, University of Guelph
Medical Image SegmentationImage AnalysisImage ProcessingMachine Learning