Secure and Explainable Fraud Detection in Finance via Hierarchical Multi-source Dataset Distillation

📅 2025-12-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address privacy leakage and lack of interpretability in collaborative financial fraud detection, this paper proposes a privacy-preserving and interpretable synthetic data generation framework. Our method introduces “tree-structured region distillation”: it converts a random forest into an axis-aligned hyperrectangular rule space and uniformly samples within it to generate a compact synthetic dataset—preserving local feature interactions while achieving full anonymization of original records. The framework simultaneously defends against membership inference attacks (attack accuracy ≈ 0.5), ensures global rule auditability, enables instance-level explanation generation, and provides calibration-aware uncertainty quantification. Evaluated on the IEEE-CIS dataset, the synthetic data achieves 85–93% compression, maintains stable AUC scores of 0.641–0.645, and boosts cross-cluster AUC to 0.687 after inter-institutional sharing.

Technology Category

Application Category

📝 Abstract
We propose an explainable, privacy-preserving dataset distillation framework for collaborative financial fraud detection. A trained random forest is converted into transparent, axis-aligned rule regions (leaf hyperrectangles), and synthetic transactions are generated by uniformly sampling within each region. This produces a compact, auditable surrogate dataset that preserves local feature interactions without exposing sensitive original records. The rule regions also support explainability: aggregated rule statistics (for example, support and lift) describe global patterns, while assigning each case to its generating region gives concise human-readable rationales and calibrated uncertainty based on tree-vote disagreement. On the IEEE-CIS fraud dataset (590k transactions across three institution-like clusters), distilled datasets reduce data volume by 85% to 93% (often under 15% of the original) while maintaining competitive precision and micro-F1, with only a modest AUC drop. Sharing and augmenting with synthesized data across institutions improves cross-cluster precision, recall, and AUC. Real vs. synthesized structure remains highly similar (over 93% by nearest-neighbor cosine analysis). Membership-inference attacks perform at chance level (about 0.50) when distinguishing training from hold-out records, suggesting low memorization risk. Removing high-uncertainty synthetic points using disagreement scores further boosts AUC (up to 0.687) and improves calibration. Sensitivity tests show weak dependence on the distillation ratio (AUC about 0.641 to 0.645 from 6% to 60%). Overall, tree-region distillation enables trustworthy, deployable fraud analytics with interpretable global rules, per-case rationales with quantified uncertainty, and strong privacy properties suitable for multi-institution settings and regulatory audit.
Problem

Research questions and friction points this paper is trying to address.

Develop a privacy-preserving dataset distillation method for financial fraud detection
Generate explainable synthetic data while maintaining model performance and privacy
Enable collaborative fraud analytics across institutions with auditable and interpretable results
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tree-region distillation creates compact, auditable synthetic datasets
Uniform sampling within rule hyperrectangles preserves local feature interactions
Explainability via aggregated rule statistics and per-case rationales
🔎 Similar Papers
No similar papers found.
Y
Yiming Qian
Institute of High Performance Computing, A*STAR, Singapore
Thorsten Neumann
Thorsten Neumann
Professor University of Applied Sciences Neu-Ulm
Empirical FinanceAsset ManagementPublic InvestmentTime-varying Parameters
X
Xueyining Huang
Xi’an Jiaotong–Liverpool University, China
D
David Hardoon
Standard Chartered Bank, Singapore
F
Fei Gao
Institute of High Performance Computing, A*STAR, Singapore
Y
Yong Liu
Institute of High Performance Computing, A*STAR, Singapore
S
Siow Mong Rick Goh
Institute of High Performance Computing, A*STAR, Singapore