A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种统一的PSMA PET/CT视觉-语言模型,用于报告生成、视觉问答和病灶分割,采用LLaVA架构并通过四阶段训练策略优化性能。
📝 Abstract
Accurate PSMA PET/CT interpretation is central to prostate cancer management, yet existing PET/CT AI models typically address isolated tasks. We propose a unified PSMA PET/CT vision-language model for report generation, visual question answering, and lesion segmentation. The framework adopts an LLaVA-style architecture, comprising a PET/CT vision encoder, an MLP-Mixer projection module, a LoRA-tuned large language model, and a 3D segmentation branch. Training followed a four-stage strategy: vision encoder pretraining, projection-layer alignment, VLM fine-tuning, and final multitask tuning. Language tasks used 5,747 PSMA PET/CT datasets with paired reports, while segmentation used the PSMA subset of AutoPET. The model outperformed PET2REP and a CT-based baseline across standard report-generation metrics, improved performance across VQA question types, and achieved higher Dice and lesion-level overlap F1 than SegAnyPET and nnUNet. These results support the feasibility of a unified framework for structured, interactive, interpretable PSMA PET/CT analysis with voxel-level grounding within a single multitask model architecture.
Problem

Research questions and friction points this paper is trying to address.

PSMA PET/CT
report generation
visual question answering
lesion segmentation
unified model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Vision-Language Model
LLaVA-style Architecture
Multitask Tuning
PSMA PET/CT
Voxel-level Grounding
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yang Xing
J. Crayton Pruitt Family Department of Biomedical Engineering, University of Florida, Gainesville, FL, USA
Jiong Wu
Jiong Wu
University of Florida
medical image analysis
S
Savas Ozdemir
Department of Radiology, University of Florida, Jacksonville, FL, USA
Yang Zhou
Yang Zhou
J. Crayton Pruitt Family Department of Biomedical Engineering, University of Florida, Gainesville, FL, USA
Boxiao Yu
Boxiao Yu
University of Florida
Deep LearningPET
Ying Zhang
Ying Zhang
Tsinghua University, Department of Industrial Engineering
Network optimizationVehicle RoutingHeuristic algorithm
Z
Zheren Zhu
Department of Radiology, University of California, San Francisco, San Francisco, CA, USA
Chenyu You
Chenyu You
Assistant Professor, Stony Brook University
Machine LearningAI for HealthComputer VisionMedical Image AnalysisMultimedia
Wei Shao
Wei Shao
Assistant Professor, University of Florida, Stanford University, University of Iowa
Medical ImagingDeep LearningImage RegistrationMedical Image AnalysisCancer Imaging
Y
Yang Lu
Division of Diagnostic Imaging, The University of Texas MD Anderson Cancer Center, Houston, TX, USA
K
Kang Wang
Department of Radiology, University of California, San Francisco, San Francisco, CA, USA
T
Tinsu Pan
Division of Diagnostic Imaging, The University of Texas MD Anderson Cancer Center, Houston, TX, USA
Y
Yang Yang
Department of Radiology, University of California, San Francisco, San Francisco, CA, USA
Kuang Gong
Kuang Gong
Assistant Professor of Biomedical Engineering, University of Florida
PETMRICTInverse ProblemMachine Learning