Institution profile

Medical University of Warsaw

Academic institutioneurope · pl
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?

Jun 10, 2026

This study addresses the susceptibility of existing multiple-choice question answering (MCQA)–based evaluations of medical large language models to guessing and answer bias, which often inflate estimates of true clinical reasoning capabilities. To mitigate these limitations, the authors introduce a more rigorous benchmark based on Polish medical licensing examinations, incorporating over 15,000 new questions spanning two additional clinical domains and implementing four structural enhancements designed to reduce inherent MCQA biases. The benchmark facilitates cross-lingual evaluation and data contamination detection, and was used to systematically assess 21 prominent large language models. Results reveal that under this more challenging setting, the top-performing model, Qwen3.5-122B, exhibits performance drops of 28.4 and 31 percentage points on the English and Polish exams, respectively, underscoring the inadequacy of standard MCQA scores as reliable indicators of genuine medical competence.

0 citationsRead paper

Conditional Fetal Brain Atlas Learning for Automatic Tissue Segmentation

Aug 06, 2025

Fetal brain MRI assessment is hindered by developmental heterogeneity, inter-site scanning protocol variations, and imprecise gestational age estimation, necessitating age-specific standardized reference atlases. To address this, we propose a conditional generative deep learning framework that integrates differentiable image registration with a conditional adversarial discriminator, enabling end-to-end mapping from gestational age to dynamic fetal brain anatomy and supporting real-time generation of continuous, age-specific brain atlases. Trained and validated on 219 normal fetal T2-weighted MRI scans, the model achieves a mean Dice score of 86.3% across six brain tissue classes, accurately recapitulating neurodevelopmental trajectories. The resulting atlas demonstrates strong robustness, cross-center generalizability, and clinically feasible inference speed (<1 second per scan). This work establishes the first learnable, updatable, and standardized reference for individualized fetal brain development assessment.

0 citationsRead paper

Semantic Mosaicing of Histo-Pathology Image Fragments using Visual Foundation Models

Aug 05, 2025

In computational pathology, large tissue sections are segmented into multiple fragments and stitched into whole-mount slide (WMS) images; however, existing boundary-based stitching methods suffer from poor robustness due to tissue defects, deformations, staining heterogeneity, and edge abrasion. To address this, we propose the first semantic-driven stitching framework leveraging Vision Foundation Models (VFMs): it extracts high-dimensional latent semantic features from pretrained pathological VFMs, constructs cross-fragment semantic correspondence candidate sets, and enables robust pose estimation and precise spatial registration. By decoupling stitching from explicit geometric boundary constraints, our method significantly improves tolerance to nonrigid deformations and staining variations. Evaluated on three public histopathological datasets, it consistently outperforms state-of-the-art methods in boundary matching accuracy, yielding more stable and higher-fidelity WMS reconstructions.

0 citationsRead paper
Recent publications

Latest Papers

Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?

Jun 10, 2026

This study addresses the susceptibility of existing multiple-choice question answering (MCQA)–based evaluations of medical large language models to guessing and answer bias, which often inflate estimates of true clinical reasoning capabilities. To mitigate these limitations, the authors introduce a more rigorous benchmark based on Polish medical licensing examinations, incorporating over 15,000 new questions spanning two additional clinical domains and implementing four structural enhancements designed to reduce inherent MCQA biases. The benchmark facilitates cross-lingual evaluation and data contamination detection, and was used to systematically assess 21 prominent large language models. Results reveal that under this more challenging setting, the top-performing model, Qwen3.5-122B, exhibits performance drops of 28.4 and 31 percentage points on the English and Polish exams, respectively, underscoring the inadequacy of standard MCQA scores as reliable indicators of genuine medical competence.

0 citationsRead paper

Conditional Fetal Brain Atlas Learning for Automatic Tissue Segmentation

Aug 06, 2025

Fetal brain MRI assessment is hindered by developmental heterogeneity, inter-site scanning protocol variations, and imprecise gestational age estimation, necessitating age-specific standardized reference atlases. To address this, we propose a conditional generative deep learning framework that integrates differentiable image registration with a conditional adversarial discriminator, enabling end-to-end mapping from gestational age to dynamic fetal brain anatomy and supporting real-time generation of continuous, age-specific brain atlases. Trained and validated on 219 normal fetal T2-weighted MRI scans, the model achieves a mean Dice score of 86.3% across six brain tissue classes, accurately recapitulating neurodevelopmental trajectories. The resulting atlas demonstrates strong robustness, cross-center generalizability, and clinically feasible inference speed (<1 second per scan). This work establishes the first learnable, updatable, and standardized reference for individualized fetal brain development assessment.

0 citationsRead paper

Semantic Mosaicing of Histo-Pathology Image Fragments using Visual Foundation Models

Aug 05, 2025

In computational pathology, large tissue sections are segmented into multiple fragments and stitched into whole-mount slide (WMS) images; however, existing boundary-based stitching methods suffer from poor robustness due to tissue defects, deformations, staining heterogeneity, and edge abrasion. To address this, we propose the first semantic-driven stitching framework leveraging Vision Foundation Models (VFMs): it extracts high-dimensional latent semantic features from pretrained pathological VFMs, constructs cross-fragment semantic correspondence candidate sets, and enables robust pose estimation and precise spatial registration. By decoupling stitching from explicit geometric boundary constraints, our method significantly improves tolerance to nonrigid deformations and staining variations. Evaluated on three public histopathological datasets, it consistently outperforms state-of-the-art methods in boundary matching accuracy, yielding more stable and higher-fidelity WMS reconstructions.

0 citationsRead paper