FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出FINSKILLOPS系统,通过控制行为维护解决SEC文件QA中的异质错误问题,利用多代理和技能补丁提高准确性。
📝 Abstract
Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter. In practice, new SEC-filing questions repeatedly expose heterogeneous errors in period, entity, evidence use, and calculation. Existing self-improvement methods can turn failures into new behaviors, but offer limited control over where a correction should apply or which previously correct answers it may break. We therefore frame post-deployment improvement as controlled behavioral maintenance: recurring failures should become scoped skill patches, and each patch should earn deployment with- out introducing regressions. We instantiate this view in FINSKILLOPS, a multi-agent system for SEC filing QA. FINSKILLOPS derives reusable skills from evidence-grounded, typed failure diagnoses and governs them through targeted validation, protected-case regression checks, negative controls, and versioned replacement or retirement. Across six financial QA benchmarks, a single frozen skill registry achieves the highest verdict-weighted correctness and reference consistency among the evaluated systems. Evolved skills raise correctness from 3.70 to 4.55 on our enhanced benchmark. In a separate 12-round operational study, only six of 33 proposed skills are promoted, while the monitoring non-correct rate falls from 20.0% to 12.5%. These results establish controlled skill scope, admission, and lifecycle management as the foundation for reliable self-improvement.
Problem

Research questions and friction points this paper is trying to address.

Financial QA
SEC filing
self-improvement
error correction
behavioral maintenance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Evolving Multi-Agent System
Controlled Behavioral Maintenance
Skill Patches
Regression Checks
Versioned Replacement
Y
Yanzhang Ma
SimpleWay.AI
Zhenghan Tai
Zhenghan Tai
University of Toronto
Information RetrievalLarge Language ModelRetrieval Augmented Generation
H
Hanwei Wu
SimpleWay.AI, McMaster University
S
Sizhe Guan
SimpleWay.AI, McMaster University
J
Jianliang Lei
SimpleWay.AI
Hailin He
Hailin He
Unknown affiliation
C
Chaolong Jiang
SimpleWay.AI
J
Jijun Chi
University of Toronto
T
Tung Sum Thomas Kwok
SimpleWay.AI, University of California, Los Angeles
B
Bohuai Xiao
SimpleWay.AI
J
Jingrui Tian
McGill University
X
Xinlu Wu
SimpleWay.AI
X
Xingao Zhan
Monash University
P
Peng Lu
Université de Montréal
Muzhi Li
Muzhi Li
The Chinese University of Hong Kong
Knowledge GraphNatural Language Processing
Yihong Wu
Yihong Wu
Université de Montréal
Machine LearningNatural Language ProcessingReinforcement Learning
Liheng Ma
Liheng Ma
PhD student, McGill University & Mila.
Geometric Deep LearningGraph Neural NetworksTime SeriesMachine Learning
S
Sicheng Lyu
SimpleWay.AI, McGill University, Mila – Quebec AI Institute
T
Tianshuo Yan
The University of Hong Kong
Junhao Zhu
Junhao Zhu
Zhejiang University
Data Lake ManagementData Integration
Y
Yaqian Xu
University of Toronto
L
Lei Ding
SimpleWay.AI, University of Manitoba
Yufei Cui
Yufei Cui
McGill University, MILA
Medical AIRAGLLM AgentPredictive Uncertainty
Ziquan Liu
Ziquan Liu
Assistant Professor, Queen Mary University of London
machine learning
B
Boyu Han
SimpleWay.AI, Stanford University