Data Quality Rule Generation with LLMs

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
The validation of data, such as customer and employee data, is an important task in many organizations. Errors in data can have severe consequences. For example, a wrong drug unit in a patient record can lead to life-threatening medication errors, and a missing street number in an address to failed deliveries. Companies often employ rule-based enterprise data quality (DQ) tools, which allow domain experts to specify rules to validate the data over time. While rule-based DQ tools are computationally efficient and provide explainable reports, maintaining a comprehensive rule set manually is challenging, as domain experts often overlook essential rules, especially in complex domains and large data volumes. Hence, closing these gaps remains an open problem in practice. In this paper, we address the challenge of automated DQ rule generation. For this, we formalize a generalizable generate-filter framework and introduce LeDQeR, an LLM-based DQ rule generation approach. First, a large language model (LLM) generates candidates rules from an observed dirty data tuple for a given rule-based DQ tool syntax. Second, we apply four filter techniques that ensure the (i) executability, (ii) correctness, and (iii) generalizability, and avoid (iv) redundancy of the generated rules. An extensive experimental evaluation suggests that LeDQeR is able to produce effective and compact rule sets for various datasets and error types.
Problem

Research questions and friction points this paper is trying to address.

Data Quality
Rule Generation
Automated DQ
Large Language Model (LLM)
Rule-based Validation
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-based DQ rule generation
generate-filter framework
executability and correctness of rules
generalizability
💼 Related Jobs
No related jobs found.
A
Anna-Christina Glock
Software Competence Center, Hagenberg, Austria
T
Thomas Hütter
Software Competence Center, Hagenberg, Austria
Johannes Fürnkranz
Johannes Fürnkranz
Johannes Kepler University, Linz
InterpretabilityPreference LearningRule LearningMultilabel ClassificationGame Playing
W
Wolfram Wöß
Johannes Kepler University, Linz, Austria
C
Christine Dominka-Kiss
Austrian Post, Vienna, Austria
Lisa Ehrlinger
Lisa Ehrlinger
Hasso Plattner Institute, University of Potsdam
Data QualityKnowledge GraphsSemantic TechnologyMetadata ManagementInformation Integration