Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque
High-performing open-source multimodal large language models (MLLMs) remain scarce for low-resource languages such as Basque. Method: This work proposes a lightweight, data-driven paradigm: autonomously constructing a high-quality Basque image–text dataset and performing end-to-end multimodal hybrid training using Llama-3.1-Instruct and the Basque language model Latxa as backbones—without Basque-specific instruction tuning. Contribution/Results: We find that only ~20% of Basque multimodal data suffices to achieve substantial performance gains, challenging the prevailing assumption that extensive language-specific supervision is required. The resulting model establishes new open-source state-of-the-art performance on Basque multimodal understanding tasks. All components—including the curated dataset, training code, and model checkpoints—are fully open-sourced. This work provides a reproducible, transferable methodology and empirical foundation for multimodal research in low-resource languages.