FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings
This work addresses the challenge of evidence provenance errors in financial question answering, which frequently arise due to similar disclosures across sections, reporting periods, and companies in SEC filings. The authors introduce the first evidence-oriented benchmark for 10-K/10-Q–based financial QA and retrieval, requiring systems to precisely align entities, reporting periods, and contextual cues to locate supporting evidence. They contribute a novel set of human-annotated hard negatives—including confounding passages—and define three distinct evaluation tasks: retrieval, re-ranking, and hard negative discrimination. Evaluated on 1,185 annotated instances, even the strongest 7B-scale model achieves only 44.8% Recall@10 using BM25, instruction-tuned embeddings, and pairwise re-ranking. Performance drops by 13.0–20.5 percentage points when hard negatives are introduced, underscoring the task’s difficulty.