Biomedical Knowledge Composition: A Software Engineering Perspective

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the persistent challenge of efficiently translating biomedical knowledge into actionable outcomes, which is hindered by technical and organizational barriers in data integration. It introduces, for the first time, a systematic engineering paradigm centered on “knowledge assembly,” drawing on software engineering principles of composability and reproducibility, with an emphasis on the construction process rather than static artifacts. The proposed approach integrates key technologies—including identifier mapping, entity disambiguation, schema alignment, evidence provenance, typed namespaces, and standardized exchange formats—to design knowledge infrastructure that supports service composition and reproducible pipelines. Through an analysis of six representative knowledge graph systems, the work identifies eight open engineering challenges, offering both domain-specific guidance and a research agenda for advancing biomedical knowledge graph development through software engineering methodologies.
📝 Abstract
Biomedical research has accumulated vast molecular, clinical, and population data, yet translating this wealth into actionable knowledge remains constrained by technical and organizational difficulties. This article presents a unified treatment of two perspectives on biomedical knowledge infrastructure. The first introduces the biomedical domain to software engineers: it explains why knowledge graphs (KGs) are the central integrative data structure in modern biomedicine, characterizes five data harmonization challenges (identifier mapping, entity resolution, schema alignment, evidence integration, and provenance tracking), surveys application domains from drug discovery to digital twins, and profiles six representative KG systems with contrasting choices. The second perspective asks why engineering biomedical knowledge infrastructure remains so difficult. We argue that a contributing root cause is limited adoption of software tooling and practices that make development in other mature domains - particularly web engineering - reliably composable and reproducible: package management, typed namespaces, canonical interchange formats, service composition protocols, reproducible pipelines, and lifecycle governance. Against this backdrop, eight open engineering challenges for biomedical data integration are catalogued, each with partial solutions but no universally adopted stack. Crucially, the article shifts emphasis from describing deployed KG instances toward the reproducible process of assembling them: reusable build pipelines, versioned dependencies, and engineering practices that let others compile and customize a KG from source rather than consuming a static artifact. Together, the two perspectives provide domain grounding for newcomers and a research agenda for software engineers seeking to make transformative contributions to biomedical knowledge infrastructure.
Problem

Research questions and friction points this paper is trying to address.

biomedical knowledge infrastructure
knowledge graphs
data integration
reproducibility
software engineering
Innovation

Methods, ideas, or system contributions that make the work stand out.

knowledge graphs
software engineering practices
reproducible pipelines
data integration
biomedical infrastructure
💼 Related Jobs
No related jobs found.
N
Natallia Kokash
Institute of Informatics, University of Amsterdam, The Netherlands
A
Adam S. Z. Bellouma
Institute of Informatics, University of Amsterdam, The Netherlands
Paola Grosso
Paola Grosso
Full Professor - University of Amsterdam
Computer NetworksFuture InternetGreen ICTe-ScienceInformation Modeling