Towards Clinically Faithful Medical Image Captioning via Enhanced Vision-Language Alignment

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决医学图像描述中的临床准确性问题,提出了一种通过增强视觉-语言对齐的框架,利用多编码器设计与辅助学习,并在推理时采用重排序方法。
📝 Abstract
Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI systems. However, unlike general image captioning, clinically reliable captioning remains challenging due to grayscale-based modalities, subtle anatomical cues, specialized medical phrasing, and variations in data quality. Despite recent advances in large vision-language models, fluent outputs do not necessarily guarantee sufficient alignment with clinical concept spaces or evaluation criteria. To address this issue, we propose a framework that strengthens clinical alignment by separating and enhancing training-time alignment and inference-time alignment. We build a medical image captioning pipeline that integrates single/dual vision encoders based on BioMedCLIP and SigLIP2, a Q-Former, and a LLaMA-based decoder, and examine the contribution of auxiliary learning for UMLS concept/type prediction. At inference, we apply single-embedding-based reranking to select the best caption among candidates, while at training we introduce MedPAIR-SCST, which combines clinically relevant rewards to shift the generative distribution toward improved clinical alignment. Our experiments show that complementary visual representations with a multi-encoder design and concept-level auxiliary learning help preserve clinically meaningful information. Furthermore, inference-time reranking provides a practical way to improve semantic and clinical alignment without additional training, whereas MedPAIR-SCST goes beyond selection by directly improving the model's distribution to generate more consistent and clinically grounded captions. These findings suggest that jointly leveraging selection-based alignment and reinforcement-learning-based alignment can promote more trustworthy medical image captioning even in data-constrained settings.
Problem

Research questions and friction points this paper is trying to address.

Medical Image Captioning
Clinical Alignment
Vision-Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Enhanced Vision-Language Alignment
MedPAIR-SCST
Inference-time Reranking
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Yunseo Lee
Yunseo Lee
University of Wisconsin-Madison, Seoul Poongnap Elementary School
FeedbackAssessmentsLearning SciencesLearning Analytics
H
Hyun Jun Kim
Department of Smart Cities, Chung-Ang University, 84 Heukseok-ro, Dongjak-gu, Seoul 06974, Republic of Korea
H
Heeseung Shin
Department of Statistics and Data Science, Chung-Ang University, 84 Heukseok-ro, Dongjak-gu, Seoul 06974, Republic of Korea
C
Changwon Lim
Department of Applied Statistics, Chung-Ang University, 84 Heukseok-ro, Dongjak-gu, Seoul 06974, Republic of Korea