Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
本文提出了一种基于参考的方法,通过比较模型变体之间的隐藏状态表示来检测偏见,解决了现有方法依赖模型输出的问题。
本文提出了一种基于参考的方法,通过比较模型变体之间的隐藏状态表示来检测偏见,解决了现有方法依赖模型输出的问题。
研究重新思考了教师-学生框架在测试时适应中的应用,通过使用不更新权重的顽固教师来解决长期稳定性问题,从而提高性能和鲁棒性。
This study addresses the challenge that existing vision-language models struggle to effectively integrate visual and textual information in Polish medical visual question answering (VQA), often over-relying on question text while neglecting image evidence. The authors present the first multi-specialty medical VQA benchmark derived from Polish physician and dentist certification exams, comprising both image-based questions and a text-only control set, along with a novel method for classifying image importance. Through systematic ablation studies—removing either images or questions—and answer-option analyses on both open-weight and commercial models, they evaluate visual grounding capabilities and reasoning biases. The best-performing model achieves 79.0% accuracy on the full test set, with only GPT-5.6 surpassing human performance on certain subsets. Models consistently underperform on image-dependent questions and can significantly exceed random guessing using answer options alone.
This work addresses the limited cultural grounding of current vision-language models, which—due to their reliance on English-centric training data—struggle to capture multimodal associations within non-English cultural contexts such as Polish. To bridge this gap, the study introduces PoVisLE, the first deep evaluation benchmark specifically designed for a single non-English culture. PoVisLE comprises 1,117 images paired with 2,366 expert-annotated visual question-answer pairs, curated under an embodied evaluation paradigm that emphasizes the interplay between language and vision within concrete cultural settings. Moving beyond superficial, template-based recognition tasks, this benchmark offers a high-quality, controllable, and challenging resource to advance the cultural adaptation and evaluation of vision-language models beyond English-dominated frameworks.
This work addresses the challenge that small language models often fail to reliably execute multi-step, dependency-rich structured graph algorithms due to error accumulation. The authors frame algorithm execution as a closed-loop prediction task, wherein the model iteratively selects operations based on the current graph state and evaluates its overall behavior through full rollbacks. Departing from conventional step-isolated evaluation, this closed-loop rollback paradigm reveals that strong single-step prediction accuracy does not necessarily ensure stable global execution. Experimental results demonstrate that suitably adapted small models can reliably perform algorithms such as traversal and coloring, yet remain vulnerable to cumulative errors in weighted graph algorithms. These findings underscore the necessity and efficacy of the proposed closed-loop evaluation framework for assessing and improving algorithmic reasoning in language models.
本文提出了一种基于参考的方法,通过比较模型变体之间的隐藏状态表示来检测偏见,解决了现有方法依赖模型输出的问题。
研究重新思考了教师-学生框架在测试时适应中的应用,通过使用不更新权重的顽固教师来解决长期稳定性问题,从而提高性能和鲁棒性。
This study addresses the challenge that existing vision-language models struggle to effectively integrate visual and textual information in Polish medical visual question answering (VQA), often over-relying on question text while neglecting image evidence. The authors present the first multi-specialty medical VQA benchmark derived from Polish physician and dentist certification exams, comprising both image-based questions and a text-only control set, along with a novel method for classifying image importance. Through systematic ablation studies—removing either images or questions—and answer-option analyses on both open-weight and commercial models, they evaluate visual grounding capabilities and reasoning biases. The best-performing model achieves 79.0% accuracy on the full test set, with only GPT-5.6 surpassing human performance on certain subsets. Models consistently underperform on image-dependent questions and can significantly exceed random guessing using answer options alone.
This work addresses the limited cultural grounding of current vision-language models, which—due to their reliance on English-centric training data—struggle to capture multimodal associations within non-English cultural contexts such as Polish. To bridge this gap, the study introduces PoVisLE, the first deep evaluation benchmark specifically designed for a single non-English culture. PoVisLE comprises 1,117 images paired with 2,366 expert-annotated visual question-answer pairs, curated under an embodied evaluation paradigm that emphasizes the interplay between language and vision within concrete cultural settings. Moving beyond superficial, template-based recognition tasks, this benchmark offers a high-quality, controllable, and challenging resource to advance the cultural adaptation and evaluation of vision-language models beyond English-dominated frameworks.
This work addresses the challenge that small language models often fail to reliably execute multi-step, dependency-rich structured graph algorithms due to error accumulation. The authors frame algorithm execution as a closed-loop prediction task, wherein the model iteratively selects operations based on the current graph state and evaluates its overall behavior through full rollbacks. Departing from conventional step-isolated evaluation, this closed-loop rollback paradigm reveals that strong single-step prediction accuracy does not necessarily ensure stable global execution. Experimental results demonstrate that suitably adapted small models can reliably perform algorithms such as traversal and coloring, yet remain vulnerable to cumulative errors in weighted graph algorithms. These findings underscore the necessity and efficacy of the proposed closed-loop evaluation framework for assessing and improving algorithmic reasoning in language models.