An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大语言模型输出有害信息的问题,提出一种基于aLoRA适配器和上下文感知路由机制的模块化校正框架,以低成本、高灵活性的方式提升模型的安全性。
📝 Abstract
Large Language Models (LLMs) are powerful zero-shot learners but remain prone to misalignment with human preferences, often producing biased, toxic, or otherwise harmful outputs. Existing alignment methods, while effective, are costly and tightly coupled to the model, limiting flexibility and scalability. We propose a modular correction framework that augments pretrained LLMs with Activated LoRA (aLoRA) adapters and a context-aware routing mechanism to eliminate harms from misaligned model responses. Our approach enables expert adapters to activate mid-sequence without invalidating the KV cache, allowing low-latency, targeted correction during generation. Each expert is trained to detect and mitigate specific harms, such as bias or toxicity. A learned router dynamically selects appropriate experts based on the models intermediate outputs. We demonstrate that our system improves alignment on standard safety benchmarks while preserving task performance, offering a lightweight and efficient path toward safer and more controllable LLM deployments.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
misalignment
harmful outputs
alignment methods
flexibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activated LoRA (aLoRA)
context-aware routing
low-latency targeted correction
dynamic expert selection
🔎 Similar Papers
No similar papers found.