Undesirable Memorization in Large Language Models: A Survey
Large language models (LLMs) exhibit “undesired memorization”—the excessive retention and leakage of sensitive training data fragments—posing critical risks for privacy violations and membership inference attacks. Method: We propose the first three-dimensional taxonomy (granularity, retrievability, extractability) to systematically characterize this phenomenon; establish a privacy–utility trade-off analysis framework unifying exposure, membership inference, and related metrics; and extend the study to emerging paradigms including retrieval-augmented generation (RAG) and diffusion language models. Through systematic literature review, empirical attribution analysis, and defense evaluation, we construct a structured knowledge graph and an open-source, dynamically updated literature repository. Results: Our work identifies six key frontiers in LLM memorization governance, advancing the field from ad hoc practice toward rigorous, systematized science.