The Realignment Problem: When Right becomes Wrong in LLMs
Current large language models (LLMs) suffer from an “alignment–reality gap”: static alignment strategies fail to adapt to dynamically evolving societal norms and policies, resulting in value misalignment, poor robustness, and high maintenance overhead. To address this, we propose TRACE—a novel framework that formalizes realignment as a programmable policy-application problem. TRACE introduces an alignment impact score to quantitatively assess preference conflicts and enables selective preference reversal, discarding, or retention—balancing correction accuracy with model performance. Leveraging a hybrid optimization pipeline—integrating conflict evaluation, preference-data categorization and filtering, and selective retraining—TRACE achieves fine-grained, low-regret updates across diverse architectures (Qwen, Gemma, Llama). Experiments demonstrate that TRACE significantly improves compliance with complex, evolving policy requirements while preserving pre-existing general-purpose capabilities.