Institution profile

NASK - National Research Institute

Academic institutioneurope · pl
Official website
Research library17linked papers
Opportunities0open roles
Selected work

Representative Papers

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

Aug 13, 2026

This study addresses the challenge that existing vision-language models struggle to effectively integrate visual and textual information in Polish medical visual question answering (VQA), often over-relying on question text while neglecting image evidence. The authors present the first multi-specialty medical VQA benchmark derived from Polish physician and dentist certification exams, comprising both image-based questions and a text-only control set, along with a novel method for classifying image importance. Through systematic ablation studies—removing either images or questions—and answer-option analyses on both open-weight and commercial models, they evaluate visual grounding capabilities and reasoning biases. The best-performing model achieves 79.0% accuracy on the full test set, with only GPT-5.6 surpassing human performance on certain subsets. Models consistently underperform on image-dependent questions and can significantly exceed random guessing using answer options alone.

0 citationsRead paper

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

Aug 07, 2026

This work addresses the limited cultural grounding of current vision-language models, which—due to their reliance on English-centric training data—struggle to capture multimodal associations within non-English cultural contexts such as Polish. To bridge this gap, the study introduces PoVisLE, the first deep evaluation benchmark specifically designed for a single non-English culture. PoVisLE comprises 1,117 images paired with 2,366 expert-annotated visual question-answer pairs, curated under an embodied evaluation paradigm that emphasizes the interplay between language and vision within concrete cultural settings. Moving beyond superficial, template-based recognition tasks, this benchmark offers a high-quality, controllable, and challenging resource to advance the cultural adaptation and evaluation of vision-language models beyond English-dominated frameworks.

0 citationsRead paper

Closed-Loop Graph Algorithm Execution with Small Language Models: Step Accuracy and Rollout Reliability

Jun 23, 2026

This work addresses the challenge that small language models often fail to reliably execute multi-step, dependency-rich structured graph algorithms due to error accumulation. The authors frame algorithm execution as a closed-loop prediction task, wherein the model iteratively selects operations based on the current graph state and evaluates its overall behavior through full rollbacks. Departing from conventional step-isolated evaluation, this closed-loop rollback paradigm reveals that strong single-step prediction accuracy does not necessarily ensure stable global execution. Experimental results demonstrate that suitably adapted small models can reliably perform algorithms such as traversal and coloring, yet remain vulnerable to cumulative errors in weighted graph algorithms. These findings underscore the necessity and efficacy of the proposed closed-loop evaluation framework for assessing and improving algorithmic reasoning in language models.

0 citationsRead paper
Recent publications

Latest Papers

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

Aug 13, 2026

This study addresses the challenge that existing vision-language models struggle to effectively integrate visual and textual information in Polish medical visual question answering (VQA), often over-relying on question text while neglecting image evidence. The authors present the first multi-specialty medical VQA benchmark derived from Polish physician and dentist certification exams, comprising both image-based questions and a text-only control set, along with a novel method for classifying image importance. Through systematic ablation studies—removing either images or questions—and answer-option analyses on both open-weight and commercial models, they evaluate visual grounding capabilities and reasoning biases. The best-performing model achieves 79.0% accuracy on the full test set, with only GPT-5.6 surpassing human performance on certain subsets. Models consistently underperform on image-dependent questions and can significantly exceed random guessing using answer options alone.

0 citationsRead paper

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

Aug 07, 2026

This work addresses the limited cultural grounding of current vision-language models, which—due to their reliance on English-centric training data—struggle to capture multimodal associations within non-English cultural contexts such as Polish. To bridge this gap, the study introduces PoVisLE, the first deep evaluation benchmark specifically designed for a single non-English culture. PoVisLE comprises 1,117 images paired with 2,366 expert-annotated visual question-answer pairs, curated under an embodied evaluation paradigm that emphasizes the interplay between language and vision within concrete cultural settings. Moving beyond superficial, template-based recognition tasks, this benchmark offers a high-quality, controllable, and challenging resource to advance the cultural adaptation and evaluation of vision-language models beyond English-dominated frameworks.

0 citationsRead paper

Closed-Loop Graph Algorithm Execution with Small Language Models: Step Accuracy and Rollout Reliability

Jun 23, 2026

This work addresses the challenge that small language models often fail to reliably execute multi-step, dependency-rich structured graph algorithms due to error accumulation. The authors frame algorithm execution as a closed-loop prediction task, wherein the model iteratively selects operations based on the current graph state and evaluates its overall behavior through full rollbacks. Departing from conventional step-isolated evaluation, this closed-loop rollback paradigm reveals that strong single-step prediction accuracy does not necessarily ensure stable global execution. Experimental results demonstrate that suitably adapted small models can reliably perform algorithms such as traversal and coloring, yet remain vulnerable to cumulative errors in weighted graph algorithms. These findings underscore the necessity and efficacy of the proposed closed-loop evaluation framework for assessing and improving algorithmic reasoning in language models.

0 citationsRead paper