Targeted Neuron Modulation via Contrastive Pair Search
Current language models lack transparency in their mechanisms for rejecting harmful requests, and prevailing intervention techniques often degrade output coherence under strong intervention intensities. This work proposes Contrastive Neuron Attribution (CNA), a method that precisely identifies, during the forward pass, the key MLP neurons responsible for distinguishing harmful from benign inputs and applies targeted modulation to steer model behavior. The study reveals, for the first time, that alignment fine-tuning transforms the inherent discriminative structures of base models into sparse, localized rejection gating mechanisms. Requiring neither gradients nor additional training, CNA enables cross-architecture analysis (e.g., Llama/Qwen) and reduces refusal rates by over 50% on standard jailbreaking benchmarks while preserving output fluency and non-degeneracy across all intervention strengths.