INDOTABVQA: A Benchmark for Cross-Lingual Table Understanding in Bahasa Indonesia Documents
This work addresses the absence of cross-lingual visual question answering (VQA) benchmarks for document tables in low-resource languages such as Indonesian by introducing the first multilingual table VQA dataset based on real-world Indonesian documents, comprising 1,593 images paired with monolingual and cross-lingual question-answer annotations. Leveraging open-source vision-language models—including Qwen2.5-VL, Gemma-3, and LLaMA-3.2—the study systematically evaluates and enhances model performance through LoRA fine-tuning and explicit incorporation of table coordinate inputs. Experimental results demonstrate that fine-tuning improves accuracy by 11.6% and 17.8% for 3B and 7B models, respectively, while integrating spatial priors yields an additional gain of 4–7%, underscoring the limitations of current models in handling complex tabular structures and low-resource linguistic settings.