Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models
This work addresses the insufficient robustness of safety mechanisms in large language models (LLMs) for low-resource languages, exemplified by Polish. We propose a lightweight, efficient multi-granularity adversarial attack that computes token importance via a small surrogate model and applies character-level perturbations guided by importance-weighted selection—requiring minimal lexical or orthographic modifications. Experiments demonstrate that our attack substantially reduces safety response rates in Polish LLMs, revealing, for the first time, their heightened vulnerability to low-cost adversarial perturbations due to scarce safety-aligned training data. To support cross-lingual robustness evaluation and defense development, we publicly release the first Polish adversarial dataset and an extensible attack framework.