On the Threat Model of Weird Generalization and Emergent Misalignment

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过调整数据集特征探讨了怪异泛化的产生条件,发现其主要受数据组成和语言影响,并对评估问题敏感。
📝 Abstract
Narrow fine-tuning on small, domain-specific datasets can produce broad and surprising changes in model behavior-a phenomenon called weird generalization (WG). Yet, it remains unclear what features of the fine-tuning data are necessary for WG to arise. Here, we address this question by investigating a range of plausibly relevant features, including dataset size, composition, language, presentation style, and novelty relative to a model's parametric knowledge. Further, since WG evaluations rely on small question sets that assess the extent of the generalization, we also analyze how sensitive this measurement is to the set of questions used. Experiments with three open-weight models on four datasets show that the degree of WG (1) depends heavily on dataset composition and language (more than on size); (2) is greater for data familiar from pretraining than for novel data; and (3) is sensitive to the set of evaluation questions used. Collectively, these results indicate that WG is a product of quite fragile properties of both training and evaluation data. As such, we argue that WG is more plausible as an adversarial threat-requiring careful data engineering-rather than as a significant hazard inherent to routine fine-tuning.
Problem

Research questions and friction points this paper is trying to address.

weird generalization
fine-tuning data
dataset composition
evaluation sensitivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Weird Generalization
dataset composition
language
evaluation questions
data engineering
🔎 Similar Papers
2024-06-17International Conference on Computational LinguisticsCitations: 6