From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过将个人文本语料库整合到小型语言模型中,利用DoRA微调方法提高多选题回答的准确性,探索了个性化认知模拟的可能性。
📝 Abstract
We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA fine-tuning. We web-crawl the search histories of 515 participants who answered 36 multiple-choice knowledge items and analyze a stratified subsample of 150 participants. For each participant, one DoRA adapter consolidates their ITC into a small language model (SLM) whose baseline correctness falls below the participants' lowest quartile. The adapter measurably writes the ITC into the weights: it fits its own participant's held-out text better than other participants' texts (dz =1.27), an individuality effect that increases with ITC size in rank order. On the generalized knowledge test, however, the adapter adds knowledge rather than alignment with the individual: log-loss match improves, whereas match accuracy under a bias-corrected PMI readout does not, and retrieval adds nothing on top. Our results demonstrate that ITCs can be consolidated into the weights of SLMs, an encouraging basis for individualized tutoring agents, and we discuss how to move from there toward a realistic simulation of episodic and semantic memory at the individual level.
Problem

Research questions and friction points this paper is trying to address.

Individual Text Corpora
Small Language Models
Episodic Memory
Semantic Memory
Multiple-Choice Question Answering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Individual Text Corpora
DoRA fine-tuning
Small Language Models
parametric individualization
🔎 Similar Papers
No similar papers found.