TopoCompress: Long Context Compression via Graph-Wired Semantic Trajectories

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为降低大型语言模型推理的成本和延迟,TopoCompress通过选择连贯的语义片段并基于混合图传播相关性分数来压缩长上下文,无需额外训练且与模型无关。
📝 Abstract
Long-context compression is essential for reducing the cost and latency of large language model inference. However, existing methods can fragment important evidence, require additional training or alignment, and often depend on the target model for effective compression. We introduce TopoCompress, a training-free and model-agnostic framework that compresses long contexts by selecting coherent semantic spans. TopoCompress first scores each span using dense and lexical query relevance together with semantic acceleration. It then constructs a hybrid graph that connects spans based on semantic similarity and sequential adjacency, and propagates the query-guided relevance scores over the graph. Across five long-context tasks-HotpotQA, 2WikiMQA, MuSiQue, Qasper, and MultiFieldQA-en-TopoCompress consistently outperforms strong compression baselines. Notably, TopoCompress achieves performance comparable to the strongest baseline while using a 4x smaller compression budget, and provides a 1.41x smaller compression time over the fastest baseline.
Problem

Research questions and friction points this paper is trying to address.

long-context compression
large language model inference
semantic spans
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free
model-agnostic
semantic trajectories
hybrid graph
query-guided relevance
🔎 Similar Papers
No similar papers found.