FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation
This work addresses the lack of a systematic evaluation benchmark for assessing the theoretical understanding and practical reasoning capabilities of large language models (LLMs) in the financial domain. To this end, we propose FIRE, a comprehensive evaluation benchmark that, for the first time, integrates questions from financial certification exams with real-world business scenarios. FIRE comprises 3,000 structured questions and open-ended problems, accompanied by standardized scoring rubrics. Leveraging a multidimensional capability taxonomy, we conduct a systematic evaluation of mainstream LLMs, including our in-house model XuanYuan 4.0. Our study not only reveals the current performance boundaries of existing models on financial tasks but also publicly releases the dataset and evaluation code, establishing a reliable benchmark to advance research in financial intelligence.