How Good LLMs Are at Answering Bangla Medical Visual Questions? Dataset and Benchmarking
This study addresses the absence of benchmark datasets for Bengali medical visual question answering (MedVQA) by introducing BanglaMedVQA—the first low-resource MedVQA dataset comprising clinically validated image-question-answer triplets in Bengali. Leveraging this dataset, the authors conduct a systematic evaluation of prominent large language models and vision-language models, including GPT-4o mini, Gemini, and Gemma-3. Experimental results reveal that current models exhibit substantially lower performance on Bengali MedVQA compared to English benchmarks, with particularly poor accuracy on complex diagnostic questions. These findings highlight a critical gap in the capability of existing multimodal models to perform fine-grained medical reasoning in low-resource languages. This work establishes essential infrastructure and empirical evidence for advancing multilingual evaluation frameworks in medical artificial intelligence.