FlexComp: One Model for Every Ratio in Context Compression

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
FlexComp通过Matryoshka式训练和两种预算选择方法解决了固定压缩比问题,实现单模型多压缩比,提高效率和准确性。
📝 Abstract
Soft context compression condenses a context into a few memory tokens that a frozen LLM consumes in place of the raw text, but existing compressors fix the compression ratio at training and inference: each deployed ratio requires a separately trained model, and the chosen ratio is applied uniformly to all inputs, whose actual needs vary drastically. We propose FlexComp, a method-agnostic framework that decouples the ratio from both training and deployment: Matryoshka-style training samples the memory budget $K$ per instance, turning one model into an any-ratio compressor, and the budget is then chosen per input by: (1) confidence-based cascade routing or (2) a lightweight learned $K$ predictor. Across ICAE, 500xCompressor, and SAC on MRQA, a single FlexComp model matches separately trained fixed-ratio specialists with minimal degradation. Cascade routing preserves over 98% of the mildest ratio's accuracy at up to 266x average compression; the $K$ predictor, in a single compression-decoding pass, reaches 158-236x within 0.7 F1 of the mildest ratio. At serving-scale batch sizes, the $K$ predictor cuts context KV cache by 50% and improves decoding throughput by 47%.
Problem

Research questions and friction points this paper is trying to address.

context compression
compression ratio
frozen LLM
memory tokens
Matryoshka-style training
Innovation

Methods, ideas, or system contributions that make the work stand out.

context compression
flexible ratio
Matryoshka-style training
confidence-based cascade routing
K predictor
🔎 Similar Papers
No similar papers found.