VLM- and LLM-Driven Multi-Agent System for PET Image Denoising

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent low signal-to-noise ratio in PET imaging and the limitations of existing denoising methods that rely on expert intervention and lack adaptability to complex scenarios. We propose a multi-agent framework driven by Vision-Language Models (VLMs) and Large Language Models (LLMs), featuring a closed-loop architecture with a rollback mechanism. This system enables dynamic image quality assessment, autonomous model selection, and decision-driven adaptive denoising workflows. Experimental results on low-dose PET data demonstrate that our approach significantly outperforms baseline methods, including UNet, GAN, and DDPM, in terms of PSNR and SSIM metrics. Consequently, this method effectively enhances both denoising performance and clinical applicability by overcoming the rigidity of conventional techniques through intelligent, autonomous optimization.
📝 Abstract
Positron emission tomography (PET) imaging suffers from limited spatial resolution and low signal-to-noise ratio, which can compromise quantitative accuracy and lesion detectability. Deep learning-based denoising methods have demonstrated strong potential for improving PET image quality. However, their practical deployment in real-world settings remains challenging, often requiring multiple specialized models and expert interventions, such as identifying motion-induced misregistration artifacts, estimating noise levels to select an appropriate denoiser, and performing lesion-focused quantitative assessment after denoising. Recent advances in vision-language models (VLMs) for image quality understanding and large language models (LLMs) for contextual reasoning provide new opportunities for automated, decision-driven workflows. Inspired by expert workflows for PET image quality enhancement, we propose an VLM- and LLM-driven multi-agent PET denoising framework that dynamically assesses image quality and lesion status, autonomously selects optimal denoising models and parameters, and enables closed-loop feedback with rollback mechanisms. Experiments were conducted on Siemens Biograph Vision Quadra PET/CT data with 1/20 and 1/50 low-dose settings. Individual module evaluations demonstrated the reliability of the agentic components, while the complete framework achieved higher PSNR and SSIM than UNet, GAN, and DDPM baselines at both dose levels. These preliminary results demonstrate the feasibility of using a closed-loop multi-agent framework to adapt PET denoising strategies to different image conditions.
Problem

Research questions and friction points this paper is trying to address.

PET image denoising
clinical deployment
expert intervention
adaptive workflow
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent System
Vision-Language Models
Large Language Models
PET Image Denoising
Closed-loop Feedback