GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of automatically generating actionable financial recommendations from enterprise operational data that integrate numerical reasoning, domain expertise, safety guarantees, and tangible business value, without relying on costly human annotations. The task is formulated as a reinforcement learning problem, and we propose Group Relative Policy Optimization (GRPO), a fine-tuning framework tailored for financial scenarios, incorporating an LLM-as-a-judge reward mechanism based on multidimensional binary scoring and a safety filter. Innovatively, we introduce CATE (Conditional Average Treatment Effect) auditing from causal inference to uncover real-world business impact that LLM-based evaluations fail to capture. Experiments demonstrate that our approach achieves a gross margin improvement of 0.0228—approximately twice that of the strongest commercial baseline—while simultaneously exhibiting the lowest downside risk and negative tail loss, all without any human-labeled data.
📝 Abstract
Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment, while avoiding recommendations that could harm the business. Direct supervision is difficult: historical decisions are not necessarily optimal, and high-quality free-form labels are expensive to obtain. We formulate financial advice generation as a reinforcement learning problem and fine-tune an open-weight language model using Group Relative Policy Optimization (GRPO). Our reward is an LLM-as-a-judge rubric that scores each recommendation across multiple binary dimensions of advice quality, augmented with a safety gate for harm prevention. Since LLM-based evaluation alone cannot confirm whether improvements reflect genuine business value rather than adaptation to the judge, we complement it with a judge-independent audit based on a standard doubly-robust Conditional Average Treatment Effect (CATE) estimator. Under this observational off-policy audit, our trained LLM achieves approximately twice the estimated gross-profit lift of the strongest evaluated commercial baseline ($0.0228$ vs.\ $0.0104$), together with the lowest downside rate and the least negative tail risk of any policy evaluated. Notably, the two evaluations do not rank the baselines identically: the untrained base model places last on the judge rubric but second on the causal audit, indicating that the audit captures a signal the judge does not. Our results demonstrate that GRPO with a finance-grounded reward signal can produce substantially more useful business recommendations than commercial LLMs, and that a judge-independent causal audit is a valuable complement to, rather than a confirmation of, LLM-as-a-judge assessment in financial NLP.
Problem

Research questions and friction points this paper is trying to address.

financial advice generation
reinforcement learning
CATE evaluation
LLM-as-a-judge
causal audit
Innovation

Methods, ideas, or system contributions that make the work stand out.

Group Relative Policy Optimization
LLM-as-a-judge
Conditional Average Treatment Effect
reinforcement learning
causal audit
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
O
Ofir Ben Shoham
Intuit
Shrutendra Harsola
Shrutendra Harsola
Senior Staff Machine Learning Scientist Lead, Intuit - India
Machine LearningLarge Language Models
V
Vignesh Subrahmaniam
Intuit
Shravan Mohan
Shravan Mohan
Intuit
Y
Yakov Gazman
Intuit
O
Oded Vainas
Intuit