ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建ReguSim环境和ReguBench基准,评估了金融合规中大语言模型代理的规则遵循情况,区分了声明理由、尝试行为、执行强制及监控证据四个因素。
📝 Abstract
LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence. We introduce ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence. In trader runs with DeepSeek V4 Pro and Gemini 3.5 Flash, visible rules reduce but do not eliminate rejected actions, and incentive or persona framing shifts behavior. A bridge study shows that trader rationales can mislead an independent monitor unless enforcement evidence is shown. In monitoring, simple structured baselines either match or exceed prompt-only LLMs. The results frame financial compliance evaluation as an audit of rule-grounded actions and evidence use, rather than a single compliance score.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
financial compliance
rule grounding
executable constraints
surveillance evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

ReguSim
financial compliance
rule grounding
monitoring benchmark
LLM agents
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.