Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Activation steering controls LLM behavior at inference time by adding learned directions to hidden states, but existing methods handle one concept at a time. Pluralistic alignment, where different stakeholders need different value emphases, requires steering multiple dimensions simultaneously. We show that naive steering produces substantial spillover: the effect intended for one value leaks into others. This parallels the treatment-versus-spillover decomposition in causal inference. We trace spillover to geometric entanglement of steering directions, captured by their Gram matrix, and derive a zero-cost correction from an activation-norm-penalized objective that decouples each direction's contribution exactly. Our end-to-end pipeline requires no fine-tuning, no reward model, and no manual prompt engineering: given only domain questions, it automatically discovers value dimensions, extracts directions, diagnoses entanglement, and applies corrected steering. On climate discourse, the correction improves the net steering effect from +5.9% to +14.0%, validated over 100,000 pairwise judgments.
Problem

Research questions and friction points this paper is trying to address.

activation steering
pluralistic alignment
spillover
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spillover-Aware
Multi-Value Steering
Activation Steering
Gram Matrix
Activation-Norm-Penalized Objective
💼 Related Jobs
No related jobs found.