Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对银行财务报表的复杂性挑战,通过创建包含999个问题的新基准数据集FinRAG-QA,并采用多阶段RAG管道优化信息检索与答案生成过程来提高金融问答系统的性能。
📝 Abstract
The comparative analysis of banks' financial statements poses significant challenges for automated question answering systems due to their complexity, substantial length, technical language, and inhomogeneity of both textual and numerical content across different jurisdictions and institutions. We introduce FinRAG-QA, a novel benchmark dataset for financial question answering, which comprises 999 practitioner-curated questions on 10 standardised indicators, grounded in 209 annual and Pillar 3 reports from 24 major European and U.S. banks spanning 2019-2023. Unlike prior financial QA benchmarks, which centre on U.S. filings and single-institution analysis, FinRAG-QA targets cross-institutional retrieval over documents averaging 198k words, longer than any existing financial QA resource. On this benchmark we evaluate a multi-stage RAG pipeline and isolate the contribution of each component. Contextual chunk enrichment combined with a retrieval-optimised embedding model raises NDCG@10 from 0.322 to 0.710; conditional on the ground truth being retrieved, a reasoning-optimised generator raises answer accuracy from 44.6% to 79.0% (+34.4 percentage points), at roughly 20x the generation latency. We further show that cross-encoder reranking degrades retrieval when the first-stage ranking is already strong, and that a single top-ranked chunk outperforms larger contexts at generation time. Experiments were run in late 2024-early 2025 with the models available at that time.
Problem

Research questions and friction points this paper is trying to address.

financial question answering
banks' financial statements
automated question answering systems
cross-institutional retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

FinRAG-QA
cross-institutional retrieval
contextual chunk enrichment
retrieval-optimised embedding model
reasoning-optimised generator
🔎 Similar Papers
2024-02-17Annual Meeting of the Association for Computational LinguisticsCitations: 26
💼 Related Jobs
No related jobs found.
Arianna Miola
Arianna Miola
Intesa Sanpaolo Innovation Center, Corso Inghilterra 3, 10138, Torino, Italy.
B
Bruno Spaccavento
Politecnico di Torino, Corso Duca degli Abruzzi 24, 10129, Turin, Italy.
L
Lorenzo Silotto
IMI CIB, Market & Counterparty Risk Management, Intesa Sanpaolo, Milan, Italy.
M
Marco Bianchetti
IMI CIB, Market & Counterparty Risk Management, Intesa Sanpaolo, Milan, Italy.
L
Luca Cagliero
Politecnico di Torino, Corso Duca degli Abruzzi 24, 10129, Turin, Italy.