RusFinChain: A Russian Benchmark for Verifiable Chain-of-Thought Reasoning in Finance with Fuzzy-Aligned Evaluation

📅 2026-07-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the absence of verifiable intermediate reasoning steps in existing Russian-language financial reasoning benchmarks, which are often limited to English or multiple-choice formats. The authors propose RusFinChain, the first Russian financial symbolic reasoning benchmark, spanning 17 domains and 172 topics, comprising 5,280 parametric samples generated via executable Python templates. Each sample includes a gold reasoning chain with intermediate numerical values, ensuring data contamination isolation and enabling automatic verification. The study introduces novel multidimensional evaluation metrics—such as fuzzy numerical alignment and soft attention alignment—that substantially improve correlation with final answer correctness (Spearman’s ρ = 0.48). Experiments on eight open-source large language models reveal a stark gap: while step-level alignment (Hard F1) reaches 0.65, answer accuracy remains around 29%, highlighting current models’ deficiencies in rigorous multi-step financial reasoning.
📝 Abstract
Multi-step symbolic reasoning is essential for robust financial analysis, yet most benchmarks neglect intermediate reasoning steps. FINCHAIN introduced verifiable Chain-of-Thought (CoT) evaluation but is limited to English. FINESSE-Bench includes a Russian block but relies on multiple-choice questions without step-level supervision. We present RusFinChain, the first Russian-language symbolic benchmark for verifiable CoT reasoning in finance. It spans 17 domains, 172 topics, and comprises 5,280 parameterized examples from executable Python templates, ensuring contamination-free evaluation. Each example includes a gold-standard reasoning chain with intermediate numeric values for automatic verification. We also introduce enhanced metrics: Fuzzy Numeric Alignment and Soft-Attention Alignment. We evaluate 8 open-weight LLMs on a stratified sample, generating 8,100 responses. Results reveal a substantial reasoning gap: models achieve Hard F1 of ~0.65 for step alignment, but only ~29% of final answers are correct. Our fuzzy and soft metrics show stronger correlation with final-answer correctness (Spearman rho approx 0.48) than the original ChainEval (rho approx 0.38-0.46), demonstrating superior diagnostic power. We release dataset, code, and evaluation framework to foster verifiable financial AI for the Russian-speaking community.
Problem

Research questions and friction points this paper is trying to address.

Chain-of-Thought Reasoning
Financial Benchmark
Russian Language
Symbolic Reasoning
Intermediate Reasoning Steps
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chain-of-Thought Reasoning
Verifiable Reasoning
Fuzzy Numeric Alignment
Financial Benchmark
Russian-language LLM Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
M. K. Arabov
Kazan Federal University, Institute of Computational Mathematics and Information Technologies, Department of Data Analysis