Institution profile

Expert.ai

Industry researcheurope · it
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Protoknowledge Shapes Behaviour of LLMs in Downstream Tasks: Memorization and Generalization with Knowledge Graphs

May 21, 2025

This work addresses how large language models (LLMs) internalize knowledge graph (KG) sequences during pretraining and generalize them into reusable knowledge. To this end, we introduce the concept of *prototype knowledge*, formally characterizing KG internalization as three structured forms: lexical, hierarchical, and topological knowledge. We propose Knowledge Activation Tasks (KATs) as a quantitative evaluation framework and establish a novel semantic-level data contamination analysis paradigm. Through controlled experiments comparing KG embedding, sequence modeling, semantic alignment, and prompting strategies, we demonstrate that prototype knowledge significantly influences Text-to-SPARQL performance, with its semantic bias strongly correlating with generalization capability. Our findings provide an interpretable, measurable empirical foundation for understanding how LLMs represent and leverage structured knowledge—bridging gaps between KG semantics and LLM pretraining dynamics.

0 citationsRead paper

SciClaims: An End-to-End Generative System for Biomedical Claim Analysis

Mar 24, 2025

Automated verification of key claims in biomedical literature faces challenges including the absence of end-to-end pipelines, error propagation in traditional NLP cascades, and poor result interpretability. To address these, this paper introduces the first LLM-native generative claim analysis system—requiring no fine-tuning—that unifies claim extraction, evidence retrieval, and truth verification via integrated prompt engineering, retrieval-augmented generation (RAG), and structured reasoning. Departing from fragile multi-stage pipelines, our approach directly generates interpretable, natural-language-supported verdicts. Evaluated across multiple benchmark tasks, it significantly outperforms existing methods, establishing a new standard for automated scientific claim analysis. The system achieves substantial gains in both accuracy and transparency, offering a novel paradigm for trustworthy scientific knowledge discovery.

0 citationsRead paper

Autoregressive Language Models for Knowledge Base Population: A case study in the space mission domain

Mar 24, 2025

Addressing the need for dynamic, real-time updates to domain-specific knowledge bases in space mission applications, this work targets efficient and accurate structured knowledge population. Method: We propose an end-to-end fine-tuned lightweight autoregressive language model that directly generates structured knowledge graph triples as JSON or SPARQL sequences. To maximize context capacity, we eliminate ontology embedding in prompts—freeing up context space for richer input/output. Domain-specific data synthesis and large-model supervised fine-tuning further enhance structured text generation capability. Contribution/Results: Experimental evaluation on the Space Mission Knowledge Base Population (KBP) task demonstrates that our compact, domain-specialized model achieves accuracy comparable to—or exceeding—that of significantly larger general-purpose LMs, while drastically reducing deployment cost and inference latency. This constitutes the first empirical validation of the effectiveness and superiority of lightweight modeling approaches for professional KBP tasks.

0 citationsRead paper

MSTS: A Multimodal Safety Test Suite for Vision-Language Models

Jan 17, 2025

This study systematically uncovers a latent safety risk in vision-language models (VLMs): hazardous behavior triggered by joint image-text inputs. Method: We introduce MSTS—the first multimodal safety evaluation suite for VLMs—comprising 400 cross-modal trigger samples spanning 40 fine-grained harm categories. We propose a novel definition and evaluation paradigm for multimodal cooperative triggering, design a cross-lingual safety assessment framework, and conduct multimodal adversarial prompting, semantic vulnerability mining, and automated classifier evaluation. Contribution/Results: We discover that non-English prompts increase harmful response rates by 3.2× on average; identify and validate “accidental safety”—a phenomenon where VLMs exhibit false compliance due to comprehension deficits; demonstrate that unimodal text-only safety testing severely underestimates real-world multimodal risks; and show that even the state-of-the-art safety classifier achieves only F1 = 0.61, underscoring the critical need for rigorous multimodal safety evaluation.

0 citationsRead paper
Recent publications

Latest Papers

Protoknowledge Shapes Behaviour of LLMs in Downstream Tasks: Memorization and Generalization with Knowledge Graphs

May 21, 2025

This work addresses how large language models (LLMs) internalize knowledge graph (KG) sequences during pretraining and generalize them into reusable knowledge. To this end, we introduce the concept of *prototype knowledge*, formally characterizing KG internalization as three structured forms: lexical, hierarchical, and topological knowledge. We propose Knowledge Activation Tasks (KATs) as a quantitative evaluation framework and establish a novel semantic-level data contamination analysis paradigm. Through controlled experiments comparing KG embedding, sequence modeling, semantic alignment, and prompting strategies, we demonstrate that prototype knowledge significantly influences Text-to-SPARQL performance, with its semantic bias strongly correlating with generalization capability. Our findings provide an interpretable, measurable empirical foundation for understanding how LLMs represent and leverage structured knowledge—bridging gaps between KG semantics and LLM pretraining dynamics.

0 citationsRead paper

SciClaims: An End-to-End Generative System for Biomedical Claim Analysis

Mar 24, 2025

Automated verification of key claims in biomedical literature faces challenges including the absence of end-to-end pipelines, error propagation in traditional NLP cascades, and poor result interpretability. To address these, this paper introduces the first LLM-native generative claim analysis system—requiring no fine-tuning—that unifies claim extraction, evidence retrieval, and truth verification via integrated prompt engineering, retrieval-augmented generation (RAG), and structured reasoning. Departing from fragile multi-stage pipelines, our approach directly generates interpretable, natural-language-supported verdicts. Evaluated across multiple benchmark tasks, it significantly outperforms existing methods, establishing a new standard for automated scientific claim analysis. The system achieves substantial gains in both accuracy and transparency, offering a novel paradigm for trustworthy scientific knowledge discovery.

0 citationsRead paper

Autoregressive Language Models for Knowledge Base Population: A case study in the space mission domain

Mar 24, 2025

Addressing the need for dynamic, real-time updates to domain-specific knowledge bases in space mission applications, this work targets efficient and accurate structured knowledge population. Method: We propose an end-to-end fine-tuned lightweight autoregressive language model that directly generates structured knowledge graph triples as JSON or SPARQL sequences. To maximize context capacity, we eliminate ontology embedding in prompts—freeing up context space for richer input/output. Domain-specific data synthesis and large-model supervised fine-tuning further enhance structured text generation capability. Contribution/Results: Experimental evaluation on the Space Mission Knowledge Base Population (KBP) task demonstrates that our compact, domain-specialized model achieves accuracy comparable to—or exceeding—that of significantly larger general-purpose LMs, while drastically reducing deployment cost and inference latency. This constitutes the first empirical validation of the effectiveness and superiority of lightweight modeling approaches for professional KBP tasks.

0 citationsRead paper

MSTS: A Multimodal Safety Test Suite for Vision-Language Models

Jan 17, 2025

This study systematically uncovers a latent safety risk in vision-language models (VLMs): hazardous behavior triggered by joint image-text inputs. Method: We introduce MSTS—the first multimodal safety evaluation suite for VLMs—comprising 400 cross-modal trigger samples spanning 40 fine-grained harm categories. We propose a novel definition and evaluation paradigm for multimodal cooperative triggering, design a cross-lingual safety assessment framework, and conduct multimodal adversarial prompting, semantic vulnerability mining, and automated classifier evaluation. Contribution/Results: We discover that non-English prompts increase harmful response rates by 3.2× on average; identify and validate “accidental safety”—a phenomenon where VLMs exhibit false compliance due to comprehension deficits; demonstrate that unimodal text-only safety testing severely underestimates real-world multimodal risks; and show that even the state-of-the-art safety classifier achieves only F1 = 0.61, underscoring the critical need for rigorous multimodal safety evaluation.

0 citationsRead paper