FinExam-10K: When Retrieval Helps Financial Reasoning?

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文构建了FinExam-10K数据集,用于评估金融推理能力,并通过对比多种模型表现,提出了一种基于FunctionGraph-RAG的方法来提高解答准确性。
📝 Abstract
Professional financial examinations require models to combine domain knowledge, calculation, and judgment, yet no benchmark covers the full CFA and FRM structure under one protocol. We introduce FinExam-10K, to our knowledge the largest reported English benchmark for this setting, with 10,198 expert-reannotated questions spanning CFA Levels I-III and FRM Parts I-II. We release 5,110 questions and sequester 5,088 for a quarterly maintained leaderboard. To separate coverage from local answerability, we report a 10,198-item Full-Coverage Track and a 7,625-item Context-Complete Reasoning Track, which is the primary basis for claims about reasoning from the supplied record. Across 17 models, the best accuracy is 85.29% overall. On the frozen Hard band, the best score is 34.68% on the Full-Coverage Track and 54.57% on the 372 context-complete items. All 17 models share 47 context-complete failures. Function-RAG and FunctionGraph-RAG rescue hundreds of errors but also overturn many correct answers, producing little or negative net gain. A gate trained only on public data decides from the question and initial response when FunctionGraph-RAG should run. On the 5,088 held-out items, the gate invokes FunctionGraph-RAG for 7.9% of questions and improves accuracy from 70.83% to 71.23% (p = .0446).
Problem

Research questions and friction points this paper is trying to address.

financial reasoning
benchmark
CFA and FRM structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

FinExam-10K
financial reasoning
benchmark
FunctionGraph-RAG
gate mechanism
🔎 Similar Papers
No similar papers found.