Individual Text Corpora Predict User-Specific Knowledge: Benchmarks of Individualized Knowledge Simulation

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用个人文本语料库和Qwen3-1.7B模型通过低秩适应方法模拟个体知识,发现模型在公开题上优于参与者但在非公开题上表现较差。
📝 Abstract
This study examines whether individual text corpora (ICs) from search histories can be used to simulate individual knowledge. We collected ICs from 316 adults, who answered 36 multiple-choice knowledge items, and compared several large language models (LLMs) on this task, of which only Qwen3-1.7B proved viable. After task-specific fine-tuning via Low-Rank Adaptation (LoRA), Qwen3-1.7B outperformed both participants and a representative German norm sample on publicly available items. On non-public questions, however, the LLM performed worse than our participants, suggesting possible training data contamination for the public questions. When integrating ICs into retrieval-augmented generation to predict individual responses, LLM-participant Match accuracies significantly exceeded chance, which demonstrates a detectable individual knowledge signal. The probabilities assigned to the participants'answers were, however, low and far below the probability of correct answers, indicating poor calibration toward individual response patterns. Knowledge-gap prediction was sub-optimal, though it improved for corpora exceeding five million tokens. We discuss our entropy based evaluation benchmarks as calibration indices for individualized knowledge simulation.
Problem

Research questions and friction points this paper is trying to address.

individual text corpora
knowledge simulation
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

individual text corpora
Low-Rank Adaptation (LoRA)
Qwen3-1.7B
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Christoph Wigbels
Bergische Universität Wuppertal
A
Ali Abusaleh
Goethe-Universität Frankfurt
M
Markus T. Jansen
Bergische Universität Wuppertal
Alexander Mehler
Alexander Mehler
Professor of Computer Science, Goethe University Frankfurt am Main
Computational HumanitiesText-technology
M
Manuel Schaaf
Goethe-Universität Frankfurt
Markus J. Hofmann
Markus J. Hofmann
Bergische Universität Wuppertal