Beyond Fluency: A Clinical Benchmark and Anomaly-Enhanced Baseline for Spine MRI Report Generation
This study addresses the critical gap in current vision-language models for lumbar MRI report generation, which often produce fluent yet clinically inaccurate diagnoses that standard natural language metrics fail to capture. To bridge this divide, the authors introduce the first clinical semantic evaluation benchmark tailored to lumbar MRI and propose an architecture-agnostic enhancement framework. This approach leverages a semi-supervised U-Net++ to generate intervertebral disc–level abnormality heatmaps, providing spatially explicit guidance to the vision-language model and thereby enhancing its anatomical sensitivity and diagnostic reliability. The method significantly improves clinical correctness while offering interpretable visual evidence, revealing for the first time the disconnect between conventional language metrics and clinical accuracy, and advancing vision-language models toward real-world clinical utility.