Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?
This study addresses the susceptibility of existing multiple-choice question answering (MCQA)–based evaluations of medical large language models to guessing and answer bias, which often inflate estimates of true clinical reasoning capabilities. To mitigate these limitations, the authors introduce a more rigorous benchmark based on Polish medical licensing examinations, incorporating over 15,000 new questions spanning two additional clinical domains and implementing four structural enhancements designed to reduce inherent MCQA biases. The benchmark facilitates cross-lingual evaluation and data contamination detection, and was used to systematically assess 21 prominent large language models. Results reveal that under this more challenging setting, the top-performing model, Qwen3.5-122B, exhibits performance drops of 28.4 and 31 percentage points on the English and Polish exams, respectively, underscoring the inadequacy of standard MCQA scores as reliable indicators of genuine medical competence.