DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization

๐Ÿ“… 2026-07-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses critical limitations in current automated chest X-ray report generationโ€”namely, deficiencies in structural coherence, anatomical completeness, and semantic fidelity. To overcome these challenges, the authors propose a novel approach that integrates supervised fine-tuning of MedGemma-4B with a clinically grounded, rule-based procedural reward mechanism. Notably, this work introduces Group Relative Policy Optimization (GRPO) to medical image report generation for the first time, enabling interpretable alignment without reliance on neural reward models. The method synergistically combines a vision-language model with multidimensional rule-based rewards encompassing structural validation, anatomical checklist adherence, semantic similarity, and length constraints. In a blinded evaluation involving 69 cases, the system achieved an impression accuracy of 27.2% and a medical terminology usage rate of 86.5%, significantly outperforming both Gemini 2.5 Flash and the MedGemma-4B baseline.
๐Ÿ“ Abstract
Medical imaging is a cornerstone of diagnostics, yet automated chest X-ray report generation struggles with structural adherence, anatomical completeness, and semantic faithfulness. We introduce DobicVLM, a vision-language model combining supervised fine-tuning on MedGemma-4B with Group Relative Policy Optimization (GRPO) and clinically-grounded programmatic rewards. Our approach uses interpretable, rule-based reward components; structural verification, anatomical checklist, semantic similarity, and length constraints to enforce clinical standards without neural reward models. Trained on 1,000 de-identified image-report pairs from a private clinical dataset (with ethics approval and compliance to local regulations), DobicVLM is evaluated via blinded expert review on 69 held-out cases. DobicVLM outperforms Gemini 2.5 Flash across the majority of criteria, achieving the highest impression accuracy (27.2%) and medical terminology (86.5%) compared to both Gemini 2.5 Flash and MedGemma 4B baselines, with minor trade-offs in completeness and referrals. This demonstrates GRPO's value for transparent alignment in resource-limited settings. Keywords: Vision-Language Models, Radiology Report Generation, Reinforcement Learning, Medical AI, GRPO
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Radiology Report Generation
Medical AI
Reinforcement Learning
GRPO
Innovation

Methods, ideas, or system contributions that make the work stand out.

Group Relative Policy Optimization
Programmatic Rewards
Clinically-Grounded Alignment
Vision-Language Models
Radiology Report Generation
T
Thanni Adewuyi
Dobic Health; University of Ibadan
A
Angelica Obayi
Dobic Health
A
Andem Aniekan
Dobic Health
S
Samuel Okoko
Dobic Health
A
Angel Ezendu
Dobic Health
E
Ephraim Usani
Dobic Health
A
Ademide Animasaun
Dobic Health
P
Philip Chibundu
Dobic Health
C
Christian Maurice
Dobic Health
M
Mary Donald Essien
Dobic Health
O
Oluwaseun Odunsi
Dobic Health
O
Oluwasegun Oguntuase
Dobic Health
A
Abiodun Adereni
Dobic Health