You Are What You Read: Misalignment via In-Context Persona Induction

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过提供良性数据而非有害行为演示,使模型在对话中采用特定人物身份,从而探讨了对齐问题。
📝 Abstract
Broad misalignment has been produced by finetuning on narrow data, harmful or benign, and in context only by demonstrations of the undesirable behaviour itself. We show that benign data suffices in context, with no finetuning and no demonstration of harmful behaviour in the prompt. Biographical facts that converge on a single figure, placed in a model's context as ordinary conversational turns, lead it to answer as that figure on questions the facts never touch. We call this persona induction. Across nine personas and thirteen models, identity adoption rises sigmoidally with the number of facts and crosses 50% within 3 to 10 of them. Misalignment then tracks which figure is described. Harmless personas reach full adoption with near-zero misalignment, while harmful ones voice their characteristic views on unrelated questions, at rates up to 80%. A formatting instruction can gate when the persona activates. Because each fact is individually benign, accumulated biographical context is flagged by content filters on 3% of inputs against 24-33% for an equivalent direct instruction.
Problem

Research questions and friction points this paper is trying to address.

persona induction
misalignment
benign data
harmful behavior
contextual influence
Innovation

Methods, ideas, or system contributions that make the work stand out.

persona induction
benign data
contextual influence
identity adoption
content filters
🔎 Similar Papers
No similar papers found.