Governing Agentic AI in FinTech

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the accountability challenges arising from deploying agent-based AI systems in finance when critical decision-making authority is granted without adequate governance, particularly due to insufficient verifiability. The authors propose a multi-layered governance framework that, for the first time, treats verifiability as a core constraint, introducing the concept of a “verifiability gap” and modeling reproducibility as an evidence-dependent governance profile. Through a three-phase empirical study comparing locally hosted large models with commercial frontier models, they systematically evaluate how temperature, top_p, top_k, and random seeds affect decision reproducibility, while also analyzing audit latency via execution logs and architectural configurations. Results show that under strict controls, local models achieve 320/320 exact reproductions and hosted models 959/960; however, architectural orchestration significantly influences outcomes, and historical decisions remain difficult to reconstruct—even with deterministic models—demonstrating that verifiability does not automatically improve with model scale.
📝 Abstract
Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes no random seed. Under the tightest controls each endpoint allows, a local model reproduced 320 of 320 executions, hosted models 319 of 320 and 959 of 960. Study 2 shows that orchestration is a latent policy layer. Architecture changes final actions, and no execution record repeated in any configuration at any scale. The frontier model reproduces its own actions more often than the local ones, its record no better, and loses a comparable share of its differentiation. Capability buys a higher starting point, not auditability. Study 3 shows two deterministic credit-model versions each reproduce their current action perfectly, yet the current cannot recover a historical one. We conceptualize reproducibility as a governance profile, not a scalar, yielding evidence-contingent delegation: authority is defensible only while retained evidence substantiates its exercise. Beyond finance, the framework extends to other high-stakes domains requiring auditability.
Problem

Research questions and friction points this paper is trying to address.

Agentic AI
Verifiability Gap
Reproducibility
Explainability
AI Governance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Verifiability Gap
Agentic AI
Reproducibility
Evidence-contingent delegation
Multilevel governance
💼 Related Jobs
No related jobs found.
H
Henry Han
Data Science and Artificial Intelligence Innovation Laboratory, Baylor University, USA