FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
Existing AI agents exhibit inconsistent performance in understanding and reasoning over real-world, complex financial documents, largely due to the absence of evaluation benchmarks that reflect industrial settings. This work proposes FinanceComplexQA—the first open-ended, generative benchmark specifically designed for complex-layout financial documents—encompassing six realistic scenarios, seven task types, and 2,026 challenging questions. The benchmark leverages a novel Finance-LaTeX SKILL pipeline to synthetically generate 2,000 professional documents and 6,000 bilingual question-answer pairs. Integrated with RAG, multi-hop reasoning, and an Agent-as-a-Judge evaluation framework, FinanceComplexQA enables systematic assessment of mainstream agents across critical dimensions such as numerical computation, summarization, and domain-specific analysis, thereby uncovering their strengths and limitations in practical financial contexts.