Institution profile

Adam Mickiewicz University

Academic institutioneurope · pl
Official website
Research library19linked papers
Opportunities0open roles
Selected work

Representative Papers

AI Grinding for Fun and Cryptanalysis

Aug 22, 2026

研究提出一种自动密码分析工作流程,通过生成、测试和优化假设来发现密码系统的缺陷。方法包括识别代数映射错误及分布差异,已验证八个已发布构造的失败。

0 citationsRead paper

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

Aug 13, 2026

This study addresses the challenge that existing vision-language models struggle to effectively integrate visual and textual information in Polish medical visual question answering (VQA), often over-relying on question text while neglecting image evidence. The authors present the first multi-specialty medical VQA benchmark derived from Polish physician and dentist certification exams, comprising both image-based questions and a text-only control set, along with a novel method for classifying image importance. Through systematic ablation studies—removing either images or questions—and answer-option analyses on both open-weight and commercial models, they evaluate visual grounding capabilities and reasoning biases. The best-performing model achieves 79.0% accuracy on the full test set, with only GPT-5.6 surpassing human performance on certain subsets. Models consistently underperform on image-dependent questions and can significantly exceed random guessing using answer options alone.

0 citationsRead paper

How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior

Aug 13, 2026

This study addresses the current lack of longitudinal empirical research evaluating whether large language models (LLMs) exacerbate the risk of AI-induced psychosis in scenarios involving the progressive escalation of delusional content. Employing a 30-day longitudinal qualitative design, the authors conducted a multidimensional analysis of 449 model-day interactions across 15 mainstream LLMs simulating the evolution of psychotic thought processes, integrating human ratings from four trained annotators with computational metrics such as entrainment and modality. The work introduces and validates four distinct LLM response trajectories: premature medicalization and disengagement, unprotected recognition, delayed unstable recognition, and delusion co-construction. Furthermore, it proposes a three-dimensional operational framework—timing of recognition, stability, and intervention accuracy—to quantify the risk of AI psychosis exacerbation, revealing that most models exhibit varying degrees of potential risk.

0 citationsRead paper
Recent publications

Latest Papers

AI Grinding for Fun and Cryptanalysis

Aug 22, 2026

研究提出一种自动密码分析工作流程,通过生成、测试和优化假设来发现密码系统的缺陷。方法包括识别代数映射错误及分布差异,已验证八个已发布构造的失败。

0 citationsRead paper

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

Aug 13, 2026

This study addresses the challenge that existing vision-language models struggle to effectively integrate visual and textual information in Polish medical visual question answering (VQA), often over-relying on question text while neglecting image evidence. The authors present the first multi-specialty medical VQA benchmark derived from Polish physician and dentist certification exams, comprising both image-based questions and a text-only control set, along with a novel method for classifying image importance. Through systematic ablation studies—removing either images or questions—and answer-option analyses on both open-weight and commercial models, they evaluate visual grounding capabilities and reasoning biases. The best-performing model achieves 79.0% accuracy on the full test set, with only GPT-5.6 surpassing human performance on certain subsets. Models consistently underperform on image-dependent questions and can significantly exceed random guessing using answer options alone.

0 citationsRead paper

How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior

Aug 13, 2026

This study addresses the current lack of longitudinal empirical research evaluating whether large language models (LLMs) exacerbate the risk of AI-induced psychosis in scenarios involving the progressive escalation of delusional content. Employing a 30-day longitudinal qualitative design, the authors conducted a multidimensional analysis of 449 model-day interactions across 15 mainstream LLMs simulating the evolution of psychotic thought processes, integrating human ratings from four trained annotators with computational metrics such as entrainment and modality. The work introduces and validates four distinct LLM response trajectories: premature medicalization and disengagement, unprotected recognition, delayed unstable recognition, and delusion co-construction. Furthermore, it proposes a three-dimensional operational framework—timing of recognition, stability, and intervention accuracy—to quantify the risk of AI psychosis exacerbation, revealing that most models exhibit varying degrees of potential risk.

0 citationsRead paper