MoNe: Modular Neural Memory for Efficient Long Context Inference

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出MoNe,一种轻量级模块化神经记忆体,通过两阶段设计实现长上下文推理,减少计算和GPU内存消耗,提高长文本处理性能。
📝 Abstract
We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads context in fixed-size segments via test-time learning of fast-weight neural memory networks with layer-localized gradient updates; at inference, the memory generates keys and values from the query tokens alone, with no context tokens re-read. This two-phase design decouples inference cost from context length, achieving $O(N)$ preprocessing and $O(1)$ query cost with peak GPU memory that does not grow with $N$. At 128K tokens, MoNe reduces both compute and peak GPU memory by approximately 80% compared to ICL with only 6.4% parameter overhead. MoNe generalizes to context lengths far beyond the backbone's native window, achieving strong performance on needle-in-a-haystack and word extraction benchmarks from RULER, where ICL degrades sharply.
Problem

Research questions and friction points this paper is trying to address.

Long-context inference
Transformer
Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

modular neural memory
test-time learning
layer-localized gradient updates
long-context inference
efficient inference
🔎 Similar Papers