Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs

๐Ÿ“… 2026-08-10
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of personalizing financial improvement strategies to enhance key performance indicators (KPIs) using non-random, multi-action accounting logs from small and medium-sized enterprises. To this end, the authors propose Covariate-Adjusted Residual Policy Learning (CAR-PL), a novel approach that implements action-level R-learning on multi-hot accounting data and incorporates observational support regularization to improve policy generalization. Evaluated on a dataset comprising 7,505 firms and 85,078 firm-month observations, CAR-PL achieves the highest point estimate (0.084) for gross profit KPI, recommends actions across all 34 categories with a more balanced distribution, and demonstrates no statistically significant difference compared to state-of-the-art methods in terms of revenue and gross profitโ€”thereby validating the effectiveness of KPI-oriented financial guidance ranking.
๐Ÿ“ Abstract
Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business changes rather than randomized recommendations. We formulate this setting as observational policy ranking: from pre-decision financial information, a policy selects one of 34 ledger-derived business-change categories for a target financial KPI. Using 85,078 company-month observations from 7,505 firms, we introduce Covariate-Adjusted Residual Policy Learning (CAR-PL), an action-wise R-learner that operates directly on multi-hot logs and regularizes selection by observational support. We compare CAR-PL with an uplift T-Learner, a conservative contextual value model, a zero-shot LLM, and non-personalized references on company-disjoint held-out firms under a shared model-assisted scoring rule. CAR-PL has the highest Gross Profit point estimate (0.084), the T-Learner has the highest Revenue point estimate (0.085), and the contextual value model has the highest Quick Ratio point estimate (0.062). CAR-PL and the T-Learner are not statistically separated on either growth KPI in matched company-clustered comparisons, while CAR-PL selects 33-34 categories and produces less concentrated selections across the catalog. Outcome-model-only scoring retains the same KPI-level point-estimate leader or top pair, and category rankings remain similar when the all-zero treatment reference is replaced by the most common training co-action pattern. These findings support objective-specific ranking of SMB financial guidance from multi-action accounting logs.
Problem

Research questions and friction points this paper is trying to address.

observational policy ranking
SMB financial guidance
multi-action accounting logs
business-change categories
financial KPI
Innovation

Methods, ideas, or system contributions that make the work stand out.

observational policy ranking
multi-action accounting logs
Covariate-Adjusted Residual Policy Learning
R-learner
SMB financial guidance
๐Ÿ”Ž Similar Papers
No similar papers found.
Shrutendra Harsola
Shrutendra Harsola
Senior Staff Machine Learning Scientist Lead, Intuit - India
Machine LearningLarge Language Models
V
Vignesh Subrahmaniam
Foresight-AI, Intuit
V
Vikas Raturi
Foresight-AI, Intuit
K
Kamalika Das
Foresight-AI, Intuit
Xiang Gao
Xiang Gao
Intuit
deep learning
K
Kratika Gupta
Foresight-AI, Intuit
Ruocheng Guo
Ruocheng Guo
Intuit AI Research
LLMsCausal MLData Mining
P
Padmaja Jonnalagedda
Foresight-AI, Intuit
A
Ananya Pramod
Foresight-AI, Intuit
S
Sricharan Kumar
Foresight-AI, Intuit