PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出PUMA,一个包含900个任务的基准测试集,旨在评估多模态模型在波兰文化和语言背景下的表现,涵盖文本、图像、音频和文档处理。
📝 Abstract
Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understanding and generation have been extensively studied, multimodal data processing capabilities, particularly in the context of cultures and languages other than English, have not yet been evaluated comprehensively. In this paper, we propose PUMA (Polish Unified Multimodal Assessment), a novel benchmark of 900 hand-crafted tasks designed to probe the limits of multimodal models in the Polish cultural and linguistic context. The dataset evaluates both cultural understanding and practical skill in processing text, images, audio, and visually rich documents. Our extensive evaluation of frontier commercial models, open-weights models, and specialized smaller systems highlights a significant performance gap. While top commercial models achieve high scores in visual question answering, most models struggle with complex audio or document understanding. We open-source our evaluation framework to advance localized multimodal AI research.
Problem

Research questions and friction points this paper is trying to address.

multimodal models
cultural understanding
Polish language
audio understanding
document processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal models
cultural understanding
Polish language
visual question answering
audio and document understanding