Attention Editing: A Versatile Framework for Cross-Architecture Attention Conversion

📅 2026-04-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing large language models in long-context and long-generation tasks, which stem from the substantial memory and bandwidth overhead of key-value (KV) cache in attention mechanisms. While more efficient attention architectures have been proposed, they are difficult to integrate into already-trained models without costly retraining. To overcome this, the authors introduce an Attention Editing framework that replaces the original attention mechanism with a learnable target module and employs a progressive distillation strategy to enable cross-architecture transfer—eliminating the need for full pretraining. This approach relaxes prior fine-grained structural constraints between source and target architectures and achieves, for the first time, a general and practical method for attention replacement in large-scale models. The framework successfully deploys MLA and GateSWA on Qwen3-8B and Qwen3-30B-A3B, significantly boosting inference efficiency on Ascend 910B clusters while preserving model performance.

Technology Category

Application Category

📝 Abstract
Key-Value (KV) cache memory and bandwidth increasingly dominate large language model inference cost in long-context and long-generation regimes. Architectures such as multi-head latent attention (MLA) and hybrid sliding-window attention (SWA) can alleviate this bound, but integrating them into existing models remains difficult. Prior methods impose fine-grained structural requirements on both source and target attention modules, which cannot meet the feasible requirement in practical deployment. We present Attention Editing, a practical framework for converting already-trained large language models (LLMs) with new attention architectures without re-pretraining from scratch. Attention editing replaces the original attention with a learnable target module and trains it using progressive distillation, consisting of (1) layer-wise teacher-forced optimization with intermediate activation supervision to prevent cold-start error accumulation, and (2) model-level distillation on next-token distributions, optionally regularized by weak feature matching. We instantiate the framework on two different target--MLA and GateSWA, a gated hybrid SWA design, and apply it to Qwen3-8B and Qwen3-30B-A3B. The resulting models maintain competitive performance while delivering substantial efficiency improvements, demonstrating that large-scale attention conversion is both feasible and robust. Notably, experiments are conducted on an Ascend 910B clusters, offering a practical training case study on domestic hardware.
Problem

Research questions and friction points this paper is trying to address.

KV cache
attention architecture
large language models
efficient inference
attention conversion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attention Editing
KV cache optimization
progressive distillation
multi-head latent attention
hybrid sliding-window attention
Z
Zhen Cheng
China Merchants Bank Artificial Intelligence Laboratory
H
Hao-Bo Yang
China Merchants Bank Artificial Intelligence Laboratory
W
Wan-Yi Huang
China Merchants Bank Artificial Intelligence Laboratory
J
Jin-Long Li
China Merchants Bank Artificial Intelligence Laboratory