heterogeneous data integration

Integrates heterogeneous datasets across domains (climate, socioeconomic, spatial transcriptomics), producing merged data products, alignment pipelines, and harmonization strategies.

heterogeneousdataintegration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.6
Aug 01, 2026Aug 01, 2026
Career
Value
No comparison yet
$193K/year
Aug 01, 2026Aug 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of effectively integrating structured data with unstructured text, a longstanding barrier in data management. It presents the first systematic argument for the necessity of textual data integration and introduces a unified framework that synergistically combines natural language processing, knowledge extraction, and traditional data integration techniques. By leveraging semantic alignment, the framework achieves deep integration between textual content and structured schemas, thereby tackling key challenges inherent in heterogeneous data integration. The work comprehensively surveys existing methodologies and outstanding issues, establishing a theoretical foundation for the emerging field of textual data integration and offering clear guidance for future research and practical implementation.

Data IntegrationHeterogeneous DataStructured Data

To address the challenge of achieving FAIR interoperability and reusability for multi-source, heterogeneous scientific data—such as the RADx COVID-19 response data—in large-scale environments, this paper proposes a general-purpose, reproducible, and extensible data harmonization framework. Our approach introduces a novel harmonization paradigm based on parameterized primitive operations and automated execution tracing. It integrates a customizable data representation model, a configurable operation library, and mechanisms for transformation logging and dependency tracking—ensuring protocol reproducibility, process auditability, and transformation reusability. Evaluated in real-world deployment within the RADx Data Hub, the framework significantly lowers the barrier to entry for domain experts, improves harmonization efficiency and transparency, and enables high-quality cross-study analyses. This work provides a scalable, transferable technical pathway for FAIR-compliant data integration across diverse scientific domains.

Addresses challenges in aligning heterogeneous data setsFacilitates data harmonization for FAIR principles complianceSupports reproducible and scalable data integration processes

Contextual Graph Embeddings: Accounting for Data Characteristics in Heterogeneous Data Integration

Nov 12, 2025
YH
Yuka Haruki
🏛️ The University of Tokyo | Infomart Corporation

In heterogeneous data integration, schema matching and entity resolution are significantly affected by domain-specific characteristics, data scale, missingness rates, and attribute overlap—yet existing graph-based methods struggle to jointly leverage structural and semantic information. Method: This paper proposes a context-aware graph embedding framework that unifies tabular structure, column-level textual descriptions, and external knowledge, employing graph neural networks for joint encoding and embedding learning of multi-source heterogeneous data. Contribution/Results: A key innovation is the context-enhancement mechanism, which systematically uncovers how data characteristics influence matching performance and empirically demonstrates that contextual modeling substantially improves robustness and accuracy—especially under challenging conditions such as high missingness rates and prevalent numeric columns. Extensive experiments across multiple domain-specific benchmark datasets show that our method consistently outperforms state-of-the-art graph-based baselines.

Addressing dataset characteristics' impact on data integration effectivenessAutomating schema matching and entity resolution in heterogeneous datasetsImproving matching reliability with contextual graph embeddings

Interactive Data Harmonization with LLM Agents

Feb 10, 2025
AS
Aécio Santos
🏛️ New York University | Federal University of Technology - Paraná

Standardizing multi-source heterogeneous clinical data remains challenging due to schema misalignment, terminological heterogeneity, and variability in data collection practices. Method: This paper proposes a novel interactive data harmonization paradigm powered by LLM-based agents, integrating domain expert knowledge with large language model reasoning. Through an interactive UI, users progressively construct harmonization pipelines, supported by core components including schema mapping, semantic alignment, and a standardized primitive library. Contribution/Results: Unlike end-to-end black-box approaches, our work introduces the first human-in-the-loop, incrementally generated, and on-demand reusable pipeline construction mechanism. Experiments demonstrate a ~70% reduction in manual coding effort, significantly improved harmonization consistency and reproducibility, and empirically validated effectiveness and generalizability on real-world clinical datasets.

Automating data harmonization processesCreating reusable data mapping pipelinesIntegrating datasets from diverse sources

This work addresses the challenges of data integration arising from heterogeneity in data schemas, value representations, and domain-specific conventions. To this end, it proposes a data harmonization framework that synergistically combines programmatic interfaces with natural language interaction. The framework integrates schema- and value-matching algorithms, AI-augmented reasoning, and composable harmonization primitives, enabling users to flexibly construct reusable harmonization pipelines via a Python API. Simultaneously, it offers domain experts a conversational interface to explore, validate, and refine results using natural language. By coupling automated matching with iterative user feedback, the system significantly enhances both the efficiency and usability of data harmonization, as demonstrated in two representative application scenarios.

data harmonizationheterogeneous dataintegrative analysis

Latest Papers

What's happening recently
View more

This work addresses the lack of end-to-end public benchmarks for data integration by introducing MaDI-Bench, the first comprehensive benchmark that spans the entire relational table integration pipeline—including schema matching, value normalization, entity blocking, entity matching, and data fusion. To mitigate benchmark saturation, MaDI-Bench incorporates a set of foundational cross-domain tasks along with a mechanism for generating extensible task variants. The benchmark supports both step-wise and end-to-end evaluation of system performance, validated through diverse pipelines ranging from manual and optimal combinations to large language model (LLM)-based approaches. All resources are publicly released to foster reproducible and holistic assessment of data integration systems.

data fusiondata integrationend-to-end benchmark

This work addresses the lack of a unified evaluation framework for knowledge graph integration pipelines, which hinders systematic comparison and selection of methods. To bridge this gap, the paper introduces KGI-Bench, the first comprehensive benchmark specifically designed for evaluating knowledge graph data integration. KGI-Bench assesses integration performance across three key dimensions—coverage, correctness, and consistency—when incorporating heterogeneous input data (structured, semi-structured, and unstructured) into a target knowledge graph. Using a curated dataset in the movie domain, the benchmark evaluates twelve representative integration pipelines, revealing significant performance variations attributable to input data types and architectural choices. The results demonstrate the effectiveness and practical utility of KGI-Bench in enabling rigorous, reproducible evaluation of knowledge graph integration approaches.

benchmarkdata integrationknowledge graph

This work proposes the first fully large language model–driven, end-to-end data integration framework that eliminates the need for manual configuration, which traditionally incurs high costs and low efficiency. The system autonomously generates a complete integration pipeline encompassing schema mapping, value normalization, entity matching, and conflict resolution without human intervention. Evaluated on three real-world domains—gaming, music, and enterprise data—the GPT-5.2–based framework achieves integration performance comparable to or surpassing that of handcrafted systems. Notably, it accomplishes this at a remarkably low cost of approximately $10 per execution, substantially reducing human labor and operational overhead.

data integrationend-to-end automationhuman effort reduction

This study addresses the limitations of traditional climate data retrieval systems, which rely on fragmented interfaces and metadata and struggle to support complex queries across multi-source simulation data for effective decision-making. For the first time, this work systematically introduces knowledge graph technology into climate services by developing a semantic knowledge graph grounded in an expert-informed ontology. The resulting framework integrates heterogeneous climate simulation datasets and enables joint semantic querying over models, variables, and spatiotemporal extents. Through comprehensive knowledge graph construction, ontology modeling, semantic integration, and an open API, the project achieves cross-dataset and cross-scenario semantic interoperability and intelligent data exploration. The authors publicly release the knowledge graph and ontology framework, significantly enhancing the efficiency of climate data discovery and its utility for informed decision support.

climate changeclimate modelsclimate services

This study addresses the challenge of achieving semantic interoperability across interdisciplinary research data, which is hindered by heterogeneous metadata schemas and ontologies. To overcome this, the authors propose DCAT-AP+, an extensible generic application profile that supports domain specialization through an upper contextual linking layer. Building on LinkML, they develop ChemDCAT-AP, a chemistry-specific sub-profile that enables unified modeling of research data and its provenance. The framework maintains compatibility with the original DCAT-AP while supporting schema inheritance, sub-profile generation, type alignment, and format transformation. Validation in the domains of chemistry and catalysis demonstrates, for the first time, that ChemDCAT-AP effectively facilitates semantic interoperability and data integration across closely related scientific disciplines.

cross-domain data integrationDCAT-APdomain ontologies

Hot Scholars

LB

Lei Bai

Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery
FL

Fenghua Ling

Shanghai Artificial Intelligence Laboratory
AI4ClimateClimate predictionWeather prediction
XX

Xiao Xiang Zhu

Technical University of Munich
Earth ObservationAI4EOSignal ProcessingData Science
KS

Konrad Schindler

Professor of Photogrammetry and Remote Sensing, ETH Zurich
PhotogrammetryRemote SensingImage AnalysisComputer Vision
YX

Yiqun Xie

University of Maryland
Spatial data miningGeoAIMachine learning