Protoknowledge Shapes Behaviour of LLMs in Downstream Tasks: Memorization and Generalization with Knowledge Graphs
This work addresses how large language models (LLMs) internalize knowledge graph (KG) sequences during pretraining and generalize them into reusable knowledge. To this end, we introduce the concept of *prototype knowledge*, formally characterizing KG internalization as three structured forms: lexical, hierarchical, and topological knowledge. We propose Knowledge Activation Tasks (KATs) as a quantitative evaluation framework and establish a novel semantic-level data contamination analysis paradigm. Through controlled experiments comparing KG embedding, sequence modeling, semantic alignment, and prompting strategies, we demonstrate that prototype knowledge significantly influences Text-to-SPARQL performance, with its semantic bias strongly correlating with generalization capability. Our findings provide an interpretable, measurable empirical foundation for understanding how LLMs represent and leverage structured knowledge—bridging gaps between KG semantics and LLM pretraining dynamics.