๐ค AI Summary
This study addresses the challenges of data sparsity and scarce winning signals in public procurement recommendation research by constructing the first large-scale knowledge graph benchmark for French public procurement. By integrating multi-source heterogeneous side information with semantic correlation analysis, the proposed framework effectively mitigates data sparsity and facilitates competition-aware recommendation modeling. The project provides comprehensive statistical analyses and fills a critical gap in publicly available datasets within this domain. Furthermore, it establishes a novel evaluation benchmark for recommendation methods in high-stakes decision-making scenarios. Ultimately, this work significantly advances both academic research and practical applications of intelligent recommendation systems in public procurement by offering a robust foundation for future studies and real-world deployment.
๐ Abstract
Public procurement represents a major economic activity, where public institutions allocate contracts to companies through competitive tendering processes. Despite its importance, this domain remains underexplored by recommender systems, largely due to the lack of publicly available datasets capturing its complexity. In this paper, we introduce TenderKG, a large-scale knowledge graph dataset constructed from French public procurement data covering the period 2021--2023. The dataset models the procurement ecosystem through heterogeneous entities, including companies, tenders, lots, and domain-specific taxonomies of work domains, connected via rich semantic and structural relations. A key specificity of this setting is that only the awarded companies are visible, resulting in sparse explicit signals of awarded interactions. To overcome this limitation, TenderKG integrates extensive side information on the actors in the French tender market and the tenders, including textual descriptions, hierarchical classifications, and geographical features, enabling the study of knowledge-aware recommendation in a highly constrained and competitive environment. We provide detailed statistics and analyses of the dataset, highlighting its structural properties, sparsity patterns, and domain-specific characteristics. We believe TenderKG opens new research directions in bidder recommendation, knowledge graph-based recommendation, competition-aware matching, and provides a valuable benchmark for evaluating methods in real-world, high-stakes decision-making scenarios.