Mirror, Mirror on the Wall: Prompt Echoing in Small Instruct Language Models

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了小型语言模型中的提示回声现象,通过分析不同模型家族发现该现象主要由模型的内部诱导机制引起而非训练数据集内容泄露。
📝 Abstract
Prompt echoing is a recognized failure mode of instruct language models, in which a model instead of generating a response, mirrors the provided prompt, even though it did not receive a specific instruction to do so. Is this phenomenon a sign of the model leaking the content of its training dataset, or is it rather caused by a misaligned behavior of the internal induction/copying mechanisms? We investigate prompt echoing small language models from different families (Gemma, Llama, Qwen, SmolLM and OLMo) and show that echoing prompts are likely to have partial overlap with the training dataset but the phenomenon is primarily driven by the model's induction heads.
Problem

Research questions and friction points this paper is trying to address.

prompt echoing
small language models
training dataset
induction heads
Innovation

Methods, ideas, or system contributions that make the work stand out.

prompt echoing
small language models
induction heads
🔎 Similar Papers
No similar papers found.
I
Inez Okulska
Centre for Credible AI, Warsaw University of Technology
B
Bartosz Naskręcki
Centre for Credible AI, Warsaw University of Technology; Adam Mickiewicz University Poznan
J
Jan Piotrowski
Centre for Credible AI, Warsaw University of Technology
Tomasz Steifer
Tomasz Steifer
Polish Academy of Sciences
machine learning & AI theory