FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming
This study addresses the inadequacy of existing safety evaluation benchmarks in capturing domain-specific financial risks—such as regulatory non-compliance, fraud inducement, and systemic trust erosion—by proposing the first two-tier threat taxonomy that integrates global financial regulatory standards (e.g., ISO/IEC 27001) with expert knowledge. Building upon this framework, the authors generate context-rich red-teaming prompt seeds derived from real-world financial documents to construct a scalable safety evaluation framework for large language models in finance. Deployed within the regulatory sandbox of the Korea Financial Security Institute, the approach employs expert-validated assessment rubrics that reduce critical false positive rates from 28% to 12%, substantially outperforming generic static rubrics and enabling high-fidelity, operationally viable AI safety evaluations in financial contexts.