Gated Differentiable Working Memory for Long-Context Language Modeling
This work addresses key challenges in long-context language modeling—attention dilution, loss of critical information, and poor generalization to novel test-time distributions—by formalizing test-time adaptation as a memory integration problem under constrained computational budgets. The authors propose an information-theoretic utility metric for context segments, coupled with a differentiable working memory module and a gated write controller, to dynamically select and integrate high-value contextual information. This approach ensures global coverage while substantially reducing gradient variance and computational overhead. Empirical results demonstrate that the method matches or exceeds state-of-the-art baselines on ZeroSCROLLS and LongBench v2 using only one-quarter of the gradient update steps, establishing a new Pareto frontier in the trade-off between efficiency and performance.