Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过识别特定群体的属性并使用基于激活工程的方法来指导大型语言模型生成针对该群体特性的解释,解决了现有方法无法有效定制解释的问题。
📝 Abstract
To effectively enable people to understand new topics, explanations should be tailored to their backgrounds and abilities. So far, prompting alone has been shown to be insufficient for creating such explanations and other computational methods are missing. Therefore, this paper investigates whether LLMs can be steered to generate explanations that are tailored to a specific group of people. To this end, we propose an approach that first identifies group-specific attributes in terms of explanatory style and knowledge of a specific target group. Building on activation engineering, it then computes attribute-based steering vectors and adds them to the internal activations of an LLM during inference to enable a fine-grained steering. In our experiments, we assess the steering effectiveness of our approach in terms of specificity and factuality of the generated explanations. Additionally, we evaluate the explanations in a study with human experts from different target groups. Compared to prompting and state-of-the-art steering baselines, our approach tailors the explanations significantly better to the target group while largely maintaining factuality.
Problem

Research questions and friction points this paper is trying to address.

Attribute-Based Activation Steering
Group-Specific Explanation Generation
Large Language Models (LLMs)
Explanatory Style
Knowledge of Target Group
Innovation

Methods, ideas, or system contributions that make the work stand out.

attribute-based steering
activation engineering
fine-grained control
group-specific explanations