Institution profile

ScaDS.AI

Academic institutioneurope · de
Official website
Research library29linked papers
Opportunities0open roles
Selected work

Representative Papers

GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base

Aug 07, 2026

Traditional knowledge bases derived from large language models often suffer from ambiguity and unreliability due to their reliance on surface-level string matching, which fails to disambiguate homonyms or consolidate synonymous expressions. This work proposes a recursive knowledge extraction framework that incorporates a context-guided entity disambiguation mechanism during construction, enabling, for the first time in LLM-derived knowledge bases, effective synonym consolidation and homonym separation. The resulting knowledge base comprises 38.4 million triples, 1.6 million canonicalized entities, 207,600 integrated relations, and 66,000 unified categories. It further integrates entity linking, relation and category clustering, a SPARQL query engine, and a natural language-to-SPARQL translation module, offering an auditable, browsable, and queryable interactive web platform alongside full public data release.

0 citationsRead paper

Example-Guided Prompting for Document-Level Text Simplification

Aug 05, 2026

Large language models (LLMs) often struggle to simultaneously preserve semantic content, ensure readability, and maintain discourse coherence in document-level text simplification when relying solely on instruction-based prompting. To address this limitation, this work proposes a retrieval-augmented, exemplar-guided prompting approach that dynamically retrieves relevant examples from a parallel simplification corpus and incorporates them into the prompt, thereby guiding the model to produce more consistent and higher-quality simplified texts without requiring task-specific fine-tuning. Evaluated systematically on the OneStopEnglish corpus, the proposed method significantly outperforms pure prompting baselines and matches or exceeds the performance of supervised and planning-based systems such as T5 and PlanSimp. Furthermore, this study provides the first empirical analysis of the varying capacities among different LLMs to effectively leverage retrieved exemplars for simplification.

0 citationsRead paper

GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models

Aug 04, 2026

This work addresses the challenge that large language models (LLMs) lack explicit entity representations, which leads to entity duplication and ambiguity when constructing knowledge bases directly. To overcome this limitation, the authors propose a native disambiguation-based approach for knowledge base construction, leveraging LLM-driven knowledge extraction coupled with real-time disambiguation of entities, relations, and categories—without relying on external resources such as Wikimedia. The method achieves explicit internal normalization and yields a large-scale knowledge base comprising over one million disambiguated entities and 38.4 million triples. By simultaneously ensuring scalability, accuracy, and cost-efficiency, this approach represents a significant departure from conventional knowledge base construction paradigms.

0 citationsRead paper

Stratified Negation in RDF Rules: A Correct Approach (Extended Version)

Jul 30, 2026

This work addresses the challenge of defining well-founded semantics for RDF rule languages—such as N3 and SHACL Rules—when default negation is introduced, a problem exacerbated by sparse triple data and blank nodes that obscure dependency structures and hinder reliable stratification. To resolve this, the paper proposes a novel chaining-based stratification method that integrates multi-step dependency analysis, integrity constraint filtering, and negation-as-failure semantics. It establishes, for the first time, a stratification criterion tailored to existential rules, effectively overcoming the difficulties posed by blank nodes and sparse graph structures. The approach guarantees a unique, minimal, and justifiable semantics for RDF rules with negation, and its feasibility and practicality are demonstrated through a prototype implementation.

0 citationsRead paper

Topic-to-Timestamp Alignment by Constrained Evidence Selection

Jun 18, 2026

This work addresses the challenge of users struggling to pinpoint specific moments in meeting discussions based solely on content. To overcome this, the paper proposes a novel approach that reframes timestamp prediction as a constrained candidate selection task. Instead of directly generating timestamps, large language models such as Mistral-7B-Instruct are guided to select the most relevant segment from a set of retrieved, timestamped meeting excerpts, thereby avoiding unsupported or invalid predictions. Integrating retrieval-augmented generation (RAG) with a constrained selection mechanism, the method demonstrates significant improvements on a dataset of 200 municipal meetings and 420 queries: Recall@5 increases from 31.9% to 50.0%, mean absolute error decreases to 761 seconds, and the number of valid outputs rises from 373 to 419, substantially enhancing both accuracy and reliability in temporal localization.

0 citationsRead paper
Recent publications

Latest Papers

GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base

Aug 07, 2026

Traditional knowledge bases derived from large language models often suffer from ambiguity and unreliability due to their reliance on surface-level string matching, which fails to disambiguate homonyms or consolidate synonymous expressions. This work proposes a recursive knowledge extraction framework that incorporates a context-guided entity disambiguation mechanism during construction, enabling, for the first time in LLM-derived knowledge bases, effective synonym consolidation and homonym separation. The resulting knowledge base comprises 38.4 million triples, 1.6 million canonicalized entities, 207,600 integrated relations, and 66,000 unified categories. It further integrates entity linking, relation and category clustering, a SPARQL query engine, and a natural language-to-SPARQL translation module, offering an auditable, browsable, and queryable interactive web platform alongside full public data release.

0 citationsRead paper

Example-Guided Prompting for Document-Level Text Simplification

Aug 05, 2026

Large language models (LLMs) often struggle to simultaneously preserve semantic content, ensure readability, and maintain discourse coherence in document-level text simplification when relying solely on instruction-based prompting. To address this limitation, this work proposes a retrieval-augmented, exemplar-guided prompting approach that dynamically retrieves relevant examples from a parallel simplification corpus and incorporates them into the prompt, thereby guiding the model to produce more consistent and higher-quality simplified texts without requiring task-specific fine-tuning. Evaluated systematically on the OneStopEnglish corpus, the proposed method significantly outperforms pure prompting baselines and matches or exceeds the performance of supervised and planning-based systems such as T5 and PlanSimp. Furthermore, this study provides the first empirical analysis of the varying capacities among different LLMs to effectively leverage retrieved exemplars for simplification.

0 citationsRead paper

GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models

Aug 04, 2026

This work addresses the challenge that large language models (LLMs) lack explicit entity representations, which leads to entity duplication and ambiguity when constructing knowledge bases directly. To overcome this limitation, the authors propose a native disambiguation-based approach for knowledge base construction, leveraging LLM-driven knowledge extraction coupled with real-time disambiguation of entities, relations, and categories—without relying on external resources such as Wikimedia. The method achieves explicit internal normalization and yields a large-scale knowledge base comprising over one million disambiguated entities and 38.4 million triples. By simultaneously ensuring scalability, accuracy, and cost-efficiency, this approach represents a significant departure from conventional knowledge base construction paradigms.

0 citationsRead paper

Stratified Negation in RDF Rules: A Correct Approach (Extended Version)

Jul 30, 2026

This work addresses the challenge of defining well-founded semantics for RDF rule languages—such as N3 and SHACL Rules—when default negation is introduced, a problem exacerbated by sparse triple data and blank nodes that obscure dependency structures and hinder reliable stratification. To resolve this, the paper proposes a novel chaining-based stratification method that integrates multi-step dependency analysis, integrity constraint filtering, and negation-as-failure semantics. It establishes, for the first time, a stratification criterion tailored to existential rules, effectively overcoming the difficulties posed by blank nodes and sparse graph structures. The approach guarantees a unique, minimal, and justifiable semantics for RDF rules with negation, and its feasibility and practicality are demonstrated through a prototype implementation.

0 citationsRead paper

Topic-to-Timestamp Alignment by Constrained Evidence Selection

Jun 18, 2026

This work addresses the challenge of users struggling to pinpoint specific moments in meeting discussions based solely on content. To overcome this, the paper proposes a novel approach that reframes timestamp prediction as a constrained candidate selection task. Instead of directly generating timestamps, large language models such as Mistral-7B-Instruct are guided to select the most relevant segment from a set of retrieved, timestamped meeting excerpts, thereby avoiding unsupported or invalid predictions. Integrating retrieval-augmented generation (RAG) with a constrained selection mechanism, the method demonstrates significant improvements on a dataset of 200 municipal meetings and 420 queries: Recall@5 increases from 31.9% to 50.0%, mean absolute error decreases to 761 seconds, and the number of valid outputs rises from 373 to 419, substantially enhancing both accuracy and reliability in temporal localization.

0 citationsRead paper