EvoOntology: A Self-Evolving Ontology Layer for Data Agents

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出EvoOntology,通过自进化本体层解决数据代理与异构数据间的交互问题,利用构建代理和自我进化循环方法优化本体。
📝 Abstract
Data agents aim to fulfill natural-language instructions over heterogeneous data, including tables, files, and databases. However, data agents face a challenging agent-data gap: heterogeneous data resides outside the agent, while the agent can access it (e.g., column names and file paths) only through generic tools. Existing approaches either let agents directly explore raw data sources or inject manually constructed semantic layers into prompts. However, neither scales well to large heterogeneous data sources nor adapts to different agent behaviors. In this paper, we introduce EvoOntology, a self-evolving ontology layer for data agents. EvoOntology encapsulates the ontology as an MCP server comprising a schema layer, a content layer, and a tool layer, enabling agents to actively query and interact with the ontology at runtime. To this end, we introduce a builder agent for autonomous ontology construction and a self-evolution loop that continuously refines the ontology through attribution-guided typed edits that are accepted only after a backbone-conditional paired evaluation. Experiments on three well-adopted data-agent benchmarks with four LLM backbones demonstrate that EvoOntology consistently outperforms strong baselines and existing semantic-layer approaches, effectively bridging the agent-data gap and enabling more effective interaction with heterogeneous data. Code: https://github.com/ruc-datalab/EvoOntology
Problem

Research questions and friction points this paper is trying to address.

data agents
heterogeneous data
agent-data gap
ontology layer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Evolving Ontology
Data Agents
MCP Server
Attribution-Guided Typed Edits
Backbone-Conditional Paired Evaluation
🔎 Similar Papers