Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how normative data, in conjunction with fine-tuning and system prompts, jointly shape the reasoning and safety behaviors of large language models in high-conflict moral dilemmas. Leveraging the Social Chemistry 101 dataset, the authors apply LoRA fine-tuning to LLaMA-3.2-11B, Qwen-3.5-9B, and Pixtral-12B, comparing models trained on norm-adhering versus norm-violating examples along fairness and deception dimensions through mixed-methods analysis. The work establishes the first traceable audit chain linking upstream normative data to downstream rationale generation, revealing that alignment behaviors emerge from the interplay of data, fine-tuning, and prompting. Specifically, fine-tuning on norm-violating data shifts model rationales toward instrumental self-interest, whereas system prompts effectively mitigate this tendency, underscoring the pivotal role of prompting in alignment mechanisms.
📝 Abstract
Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather than neutral moral knowledge. We propose treating the AI system as a proxy actor and test whether dataset-level norms can shift it away from its baseline safety behavior when it faces high-conflict dilemmas. We make three contributions. First, we demonstrate in controlled experiments that norm-breaking fine-tuning yields norm-divergent actions justified by self-interested rationales, suggesting a systematic shift in patterns of justification. Second, we establish a practical audit trail linking downstream justifications to upstream norms using mixed methods. Third, we show that system prompts can both suppress and elicit these patterns. We conducted experiments on three models (LLaMA-3.2-11B, Qwen-3.5-9B, and Pixtral-12B) using Low-Rank Adaptation (LoRA) fine-tuning on Social Chemistry 101 Fairness/Cheating (norm-following vs. norm-breaking) with prompt steering. Across all three models, we find that norm-breaking fine-tuning shifts the model's default rationale style from safety compliance to instrumental self-interest, whereas system prompts can override this behavior. Our results support a distributed view of alignment in which observed behavior depends jointly on training data, fine-tuning, and prompting, motivating norm-aware documentation and rationale logging for contestable oversight.
Problem

Research questions and friction points this paper is trying to address.

normative datasets
AI alignment
model rationales
fine-tuning effects
prompt influence
Innovation

Methods, ideas, or system contributions that make the work stand out.

norm-aware alignment
rationale auditing
LoRA fine-tuning
prompt steering
instrumental self-interest
🔎 Similar Papers
No similar papers found.