Are LLMs Truly Multilingual? Exploring Zero-Shot Multilingual Capability of LLMs for Information Retrieval: An Italian Healthcare Use Case
Prior work lacks localized, clinical empirical evaluation of open-source multilingual large language models (LLMs) for comorbidity extraction from Italian electronic health records (EHRs) in zero-shot settings. Method: We systematically assess the real-time zero-shot comorbidity extraction capability of multilingual LLMs on authentic Italian clinical texts, deploying models locally and benchmarking them against rule-based matching and human annotation. Contribution/Results: Our quantitative analysis reveals that state-of-the-art multilingual LLMs underperform rule-based methods significantly and exhibit poor generalization across disease categories—highlighting fundamental limitations in domain-specific clinical text understanding. This study provides critical empirical evidence and methodological guidance on the applicability of multilingual LLMs to low-resource clinical natural language processing tasks, particularly in under-resourced languages and specialized medical domains.