Can Open-Weight Models Compete on Financial Text Comprehension?

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation of open-source large language models on real-world financial text understanding tasks by extending the Financial Touchstone benchmark, which now comprises 2,967 question-answer triplets derived from 495 international annual reports. The authors conduct a comprehensive assessment of 20 prominent models, including GLM-4.7, GLM-5, Kimi K2.6, DeepSeek V3.2, and leading closed-source counterparts. Notably, they find that certain non-reasoning open-source models—such as Kimi K2.6 and GLM-5—can match or even surpass some closed-source models in financial comprehension. The study also reveals that Chinese-language models often unjustifiably refuse legitimate queries due to content filtering mechanisms, with refusal behavior influenced by access pathways. Among evaluated models, Claude Opus 4.6 achieves the highest accuracy (88.4%), while Gemini 2.5 Pro exhibits the lowest hallucination rate (0.08%). Information retrieval errors account for 48.9% of failures. The full dataset and evaluation framework are publicly released.
📝 Abstract
Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability on real-world financial tasks remains largely untested. We updated the Financial Touchstone benchmark, which now has 2,967 question context-answer triplets across 495 international annual reports. We also apply a new set of models on the benchmark, expanding coverage from eleven to twenty models across ten providers, including recent open-weight models such as GLM 4.7, GLM 5, Kimi K2.6, and DeepSeek V3.2, as well as Alibaba's proprietary flagship Qwen3-Max. Anthropic's Claude Opus 4.6 achieves the highest accuracy (88.4%), while Google's Gemini 2.5 Pro maintains the lowest hallucination rate (0.08%). Notably, the open-weight Kimi K2.6 ranks third in accuracy, and the non-reasoning models GLM 5 and Mistral 3 rank fourth and fifth, challenging the assumption that reasoning architectures or proprietary weights are a prerequisite for strong financial comprehension. Information retrieval remains the primary bottleneck, accounting for 48.9% of all failures. We also document a new finding: geopolitical content filters in Chinese models refuse legitimate financial questions (0.08% of attempts), sometimes without clear reason, and the refusal behavior depends on the access route as much as on the model. The complete dataset and evaluation framework are publicly available.
Problem

Research questions and friction points this paper is trying to address.

financial text comprehension
open-weight models
information retrieval
hallucination
geopolitical content filtering
Innovation

Methods, ideas, or system contributions that make the work stand out.

open-weight models
financial text comprehension
benchmark evaluation
hallucination rate
geopolitical content filtering
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jan Spörer
University of St. Gallen, Switzerland