EU-Agent-Bench: Measuring Illegal Behavior of LLM Agents Under EU Law
Large language model (LLM) agents operating within the EU risk non-compliance with stringent legal frameworks such as the GDPR and the Equal Treatment Directive, yet no benchmark exists to systematically evaluate their legal adherence. Method: We introduce EU-LegalBench—the first evaluation benchmark explicitly designed for legal compliance assessment of LLM agents under EU law. It comprises (1) a manually curated, verifiable test suite covering high-risk domains including data protection, algorithmic discrimination, and research integrity, grounded in primary EU legislation; (2) a legislative citation alignment mechanism that explicitly links model behaviors to specific statutory provisions; and (3) controlled system-prompting experiments to quantify how embedding legal text affects compliance performance. Contribution/Results: Empirical results demonstrate that explicit legal prompting significantly improves compliance rates. EU-LegalBench includes a publicly available preview set and a controlled private test set, establishing a reproducible, attributable, and legally grounded paradigm for assessing LLMs’ legal safety in regulated environments.