Phare: A Safety Probe for Large Language Models
Existing LLM safety evaluations predominantly focus on performance ranking, neglecting systematic failure-mode identification. This paper introduces Phare—the first multilingual safety assessment framework centered on failure-mode diagnosis—systematically uncovering model vulnerabilities across three dimensions: hallucination and reliability, social bias, and harmful content generation. Phare innovatively integrates adversarial prompt engineering, controllable behavioral sampling, and fine-grained human annotation to identify concrete risk patterns—including sycophancy, prompt sensitivity, and stereotype reiteration—for the first time. Evaluated on 17 mainstream models, it reveals cross-lingual and cross-architectural common vulnerabilities. Furthermore, Phare provides actionable, model-agnostic improvement pathways. By shifting evaluation from aggregate scoring to diagnostic analysis, it advances the development of more robust, aligned, and trustworthy language systems.