π€ AI Summary
This study addresses the critical challenge of identifying economically meaningful comparable companies in opaque and sparsely traded private marketsβa key issue for valuation, due diligence, and risk management. The authors propose a supervised similarity learning framework based on ensemble trees, which uniquely anchors similarity on market valuations. Leveraging CatBoost, the model learns nonlinear valuation drivers from data on 270,000 global private firms, including 53,000 with observed valuation records. An interpretable similarity metric is constructed through importance-weighted co-occurrence of leaf nodes, effectively handling mixed data types and extensive missing values while preserving case-level interpretability. Empirical results demonstrate that this approach significantly outperforms conventional distance metrics and text-embedding methods in cross-industry k-nearest neighbor valuation tasks.
π Abstract
As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, identifying economically meaningful peer companies for comparison is a fundamental challenge for valuation, due diligence, portfolio construction, and risk management. We propose an ensemble tree-based supervised similarity learning framework that defines company similarity through the lens of market valuation rather than static feature matching or semantic descriptions. Specifically, we train a CatBoost gradient-boosted decision tree model on observed private company valuations and derive a valuation-aware similarity metric from importance-weighted leaf-node co-occurrences across the ensemble. The similarity metric captures shared valuation drivers while accommodating nonlinear relationships, mixed data types, and pervasive missing data common in private markets. Using a global private-market universe of approximately 270,000 companies, including more than 53,000 firms with observed or derivable post-money valuations spanning multiple industries, geographies, and deal stages, we demonstrate that the proposed similarity framework improves upon traditional distance-based and text-embedding-based approaches in downstream k-nearest-neighbor valuation tasks in the evaluated industry groups, while retaining case-based explainability.