Order in Desbordante: Techniques for Efficient Implementation of Order Dependency Discovery Algorithms

📅 2026-07-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the critical role of order dependencies (ODs) in query optimization and data cleaning, where existing discovery algorithms suffer from efficiency and scalability limitations. The study presents the first systematic analysis of engineering bottlenecks in OD discovery and introduces a high-performance implementation by re-engineering the FASTOD and ORDER algorithms in C++. These optimized algorithms are integrated into the Desbordante data profiling framework, leveraging advanced memory management and computational strategies. Compared to the original implementations, the proposed approach achieves up to a 10× speedup and reduces memory consumption by as much as 2.9×, substantially enhancing the practicality and scalability of OD discovery for real-world applications.
📝 Abstract
Science-intensive data profiling focuses on discovery and validation of various patterns in datasets. This study considers discovery of one such pattern - order dependency (OD). Simply put, OD states that some list of columns is ordered according to another one. It is of use for database query optimization, data cleaning and deduplication, anomaly detection, and much more. Existing discovery methods have approached this problem solely from the algorithmic standpoint, without focusing on the implementation side. At the same time, this problem is very computationally intensive, and therefore this part should not be ignored, as it brings ODs closer to industrial use. In this paper, we study two algorithms for OD discovery which target different OD axiomatizations - FASTOD and ORDER. We start by reimplementing these algorithms in C++ in order to speed them up and lower their memory consumption. We then analyze their bottlenecks and propose several techniques which improve their performance even further. To perform evaluation, we have implemented these algorithms inside Desbordante - a science-intensive, high-performance, and open-source data profiling tool developed in C++. Experiments have demonstrated a performance improvement of up to 3x obtained by reimplemented versions, and, with the application of our techniques, up to 10x. Memory consumption has been lowered by up to 2.9x.
Problem

Research questions and friction points this paper is trying to address.

order dependency
data profiling
algorithm implementation
computational efficiency
memory consumption
Innovation

Methods, ideas, or system contributions that make the work stand out.

order dependency
algorithm optimization
data profiling
performance acceleration
Desbordante
Y
Yakov Kuzin
Saint-Petersburg University
D
Dmitriy Shcheka
Saint-Petersburg University
M
Michael Polyntsov
Saint-Petersburg University
K
Kirill Stupakov
Saint-Petersburg University
M
Mikhail Firsov
Saint-Petersburg University
G
George Chernishev
Saint-Petersburg University