OmniKVQuant: KV Cache Quantization for Omni-LLMs

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对Omni-LLMs中KV缓存内存成本增加的问题,通过提出OmniKVQuant方法解决多模态缓存中的关键漂移和异构值几何问题。
📝 Abstract
As Omni-modal large language models (Omni-LLMs) take in audio, video and text together, their KV cache memory cost grows. KV cache quantization is the de facto approach in text-only LLMs, but its application to Omni-LLMs remains unexplored. In this paper, we analyze how TurboQuant, a representative rotation-based KV cache quantization method, behaves on multimodal caches and identify two critical issues: temporal key drift and heterogeneous value geometry. To address these, we propose OmniKVQuant, a training-free framework that (i) sets the key quantization range over each short window of the input stream; and (ii) rotates values separately per modality. On Qwen2.5-Omni and Qwen3-Omni, OmniKVQuant enables 2-bit KV caches while substantially preserving performance across seven audio-visual benchmarks. We further provide a fused Triton decode kernel that unpacks the 2-bit cache during attention, so no dense FP16 cache is ever built. Code: https://github.com/kaistmm/OmniKVQuant
Problem

Research questions and friction points this paper is trying to address.

Omni-LLMs
KV cache quantization
multimodal caches
Innovation

Methods, ideas, or system contributions that make the work stand out.

Omni-LLMs
KV Cache Quantization
Temporal Key Drift
Heterogeneous Value Geometry
Training-free Framework
🔎 Similar Papers