JustLLMGRPO: Radiographic Control for Chest X-Ray Generation

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing text-guided chest X-ray generation methods, which overly rely on optimizing image generators while neglecting the critical role of prompt formulation in generation quality. Under the constraint of a frozen Sana generator pre-adapted to chest X-rays, this study introduces Group Relative Policy Optimization (GRPO) to optimize prompt strategies for large language models, enhancing image realism and prompt alignment solely through refined input prompts. Experiments demonstrate that the proposed approach reduces the RadDINO-FID score to 26.780 on CheXGenBench—a 50.6% improvement over the baseline—while preserving source prompt alignment (0.696 vs. 0.695). Furthermore, it achieves state-of-the-art performance in distribution coverage and downstream classification utility, revealing substantial untapped potential in prompt expression for medical image synthesis.
📝 Abstract
Text-conditioned chest X-ray generation aims to synthesize realistic radiographs that faithfully depict specified findings. Existing work has primarily improved quality by updating image generators, implicitly treating prompts as fixed after CXR-domain adaptation. We show that this generator-centric view leaves a substantial optimization dimension underexplored. With a CXR-adapted Sana generator frozen, one-pass reformulation by an unmodified LLM reduces RadDINO-FID from 54.225 to 27.572. Prompt analysis shows that the LLM suppresses temporal comparisons, uncertainty, and other non-renderable report content while emphasizing visible radiographic findings. However, unconstrained reformulation reduces BioViL-T alignment with source prompts from 0.695 to 0.609. We therefore introduce JustLLMGRPO, which applies standard Group Relative Policy Optimization (GRPO) only to the LLM prompt policy while keeping Sana frozen. Group-relative radiology-aware image feedback retains visual focus while preserving source-prompt alignment. On CheXGenBench, JustLLMGRPO reduces RadDINO-FID to 26.780, a 50.6% improvement over direct prompting, while maintaining alignment (0.696 versus 0.695). It also achieves state-of-the-art distribution coverage and downstream classification utility. These results show that substantial performance can remain latent in how radiographic information is expressed to an adapted generator. Code is publicly available at https://github.com/pxcai/JustLLMGRPO.
Problem

Research questions and friction points this paper is trying to address.

chest X-ray generation
text-conditioned synthesis
prompt reformulation
radiographic control
image-text alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prompt Reformulation
Group Relative Policy Optimization
Text-to-X-ray Generation
Radiology-aware Feedback
Frozen Generator
🔎 Similar Papers
No similar papers found.
P
Pengxiang Cai
The Hong Kong University of Science and Technology (Guangzhou)
Xiaohan Li
Xiaohan Li
Walmart Inc.
Data MiningRecommender systemMedical AI
A
Anglin Liu
The Hong Kong University of Science and Technology (Guangzhou)
Qingyuan Zeng
Qingyuan Zeng
Xiamen University
computer vision
Z
Zexun Li
The Hong Kong University of Science and Technology (Guangzhou)
Jintai Chen
Jintai Chen
Assistant Professor@HKUST(GZ)
AI for HealthcareMultimodal LearningDeep Tabular Learning