What Attention Recalls and Recurrence Controls in Hybrid Language Models

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过两种干预方法探讨了混合语言模型中注意力和循环状态的角色,发现注意力负责精确检索,而循环状态控制输出语言和个性。
📝 Abstract
Hybrid language models combine attention with a fixed-size recurrent state, but the role of each channel remains unclear. We introduce two cache-level interventions. Split-prefill keeps only the KV cache or only the recurrent state from a prefilled context, then generates an answer. State-swap pairs the KV cache from one context with the recurrent state from another in a single forward pass. On Qwen3.5 and Falcon-H1, the two channels split sharply by function. Exact retrieval survives only through attention (64-98% of full accuracy) and collapses to zero through recurrence. Output language and persona reverse the pattern: both survive recurrence (70-80% and 3-5x) while KV-only drops to ~1% language accuracy. State-swap confirms this causally: the answer takes its value from the KV side and its language from the recurrent side. Recurrent-only generation also accepts words that were never in the context but share meaning or parts with seen items. Attention provides a lookup over what was said; the recurrent state shapes how the model says it next.
Problem

Research questions and friction points this paper is trying to address.

attention
recurrence
hybrid language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

attention
recurrence
hybrid language models
KV cache
retrieval
🔎 Similar Papers