CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses evidence trust bias, verification latency, and insufficient hallucination defense in cross-modal financial question answering by proposing a nine-agent collaborative framework. By constructing a typed claim ledger, an asymmetric evidence authority mechanism, and an adaptive debate loop, the system achieves claim-level dynamic auditing and fine-grained verification. Experimental results demonstrate that this approach attains a faithfulness score of 0.889, significantly outperforming baselines such as Graph-RAG. Furthermore, it exhibits active refusal capabilities under conditions of insufficient evidence, thereby effectively enhancing the reliability and safety of financial question answering systems.
📝 Abstract
Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagreement, debate verifies an aggregate report rather than individual claims, and such verification occurs only after drafting, leaving inter-agent errors undetected until the final text. To close this gap, we present CLAIR-Fin, a nine-agent framework that decomposes each question into atomic claims maintained in a typed Financial Claim Ledger. Each claim is resolved through Asymmetric Evidence Authority, which conditions evidence trust on claim type rather than treating all modalities as equally reliable; Chain-of-Custody Verification, which checks grounding at the hand-off between drafting and adversarial review rather than only at the pipeline's exit; an Adaptive Rebuttal Cycle, which routes contested claims through adversarial debate whose depth scales with what that debate finds; and a terminal entailment audit paired with a continuous Hallucination Risk Index that distinguishes claims that passed scrutiny from claims never contested. We evaluate CLAIR-Fin on BB-FinQA-X, a 500-question cross-modal financial evaluation set built from Bangladesh Bank Annual Report material, stratified by query type, format, and difficulty. Relative to a single-pass retrieval-augmented generation baseline, it raises faithfulness ($0.780 \rightarrow 0.889$) while abstaining on 5.4% of questions when evidence is insufficient rather than forcing an unsupported response, and it exceeds stronger retrieval-strategy baselines such as HyDE and Graph-RAG on faithfulness ($\leq 0.874$).
Problem

Research questions and friction points this paper is trying to address.

Hallucination
Multi-Agent Systems
Cross-Modal Financial QA
Claim-Level Verification
Retrieval-Augmented Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Multi-Agent Framework
Asymmetric Evidence Authority
Chain-of-Custody Verification
Adaptive Rebuttal Cycle
Hallucination Risk Index
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.