VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference
为解决视觉大语言模型长上下文推理的高计算和内存开销问题,提出VisCache框架,通过选择性前向关键帧和非对称KV剪枝方法提升效率。
为解决视觉大语言模型长上下文推理的高计算和内存开销问题,提出VisCache框架,通过选择性前向关键帧和非对称KV剪枝方法提升效率。
This work addresses the limitations of centralized prediction-based routing in large language models (LLMs), which often leads to misaligned information risks and scalability bottlenecks. To overcome these issues, the paper introduces, for the first time, a reverse auction mechanism into LLM routing and proposes the error-aware EA-RAM framework. In this framework, model providers autonomously bid their success rates and costs, while explicit modeling of dual sources of noise—arising from both prediction and evaluation—enables robust handling of uncertainty. The mechanism is shown to be Bayesian incentive-compatible and individually rational, with a provable upper bound on social welfare loss. Empirical results demonstrate that EA-RAM consistently outperforms centralized baselines across both simulated and real-world benchmarks, maintaining robustness under dual-error conditions and significantly advancing the cost-performance Pareto frontier.
To address the high computational cost and difficulty in balancing convergence efficiency of thresholding-based Newton-type methods for sparse optimization, this paper proposes a Compressed Newton Thresholding Optimization framework. It compresses the Newton direction onto a low-dimensional subspace and incorporates diagonal regularization, yielding two novel algorithms: Compressed Newton Hard Thresholding Pursuit (CNHTP) and Compressed Newton Optimal Thresholding Pursuit (CNOTP). Under the Restricted Isometry Property (RIP), we establish rigorous theoretical guarantees of global convergence. The algorithms integrate compressed direction computation, adaptive thresholding selection, and regularization, preserving Newton-type convergence rates while significantly reducing per-iteration complexity. Experiments demonstrate that the proposed methods match state-of-the-art algorithms in recovery success rate and solution accuracy, while achieving superior computational efficiency and robustness.
This work addresses the sparse solution recovery problem for linear systems involving concatenated orthogonal matrices. To exploit their structural properties, we propose a splitting-based alternating optimization framework—supporting both two-block and multi-block decompositions—that relies solely on matrix-vector products and low-dimensional orthogonal projections, thereby avoiding explicit matrix inversion or storage. By decomposing the large-scale system into coupled subsystems and solving them cooperatively via iterative updates, the algorithm is proven to converge globally to the sparse solution under a coherence constraint. Compared to mainstream iterative methods—including Orthogonal Matching Pursuit (OMP) and Iterative Shrinkage-Thresholding Algorithm (ISTA)—the proposed approach significantly reduces iteration counts while achieving faster convergence and enhanced numerical stability. This yields an efficient, scalable, and structurally aware paradigm for high-dimensional sparse signal recovery.
为解决视觉大语言模型长上下文推理的高计算和内存开销问题,提出VisCache框架,通过选择性前向关键帧和非对称KV剪枝方法提升效率。
This work addresses the limitations of centralized prediction-based routing in large language models (LLMs), which often leads to misaligned information risks and scalability bottlenecks. To overcome these issues, the paper introduces, for the first time, a reverse auction mechanism into LLM routing and proposes the error-aware EA-RAM framework. In this framework, model providers autonomously bid their success rates and costs, while explicit modeling of dual sources of noise—arising from both prediction and evaluation—enables robust handling of uncertainty. The mechanism is shown to be Bayesian incentive-compatible and individually rational, with a provable upper bound on social welfare loss. Empirical results demonstrate that EA-RAM consistently outperforms centralized baselines across both simulated and real-world benchmarks, maintaining robustness under dual-error conditions and significantly advancing the cost-performance Pareto frontier.
To address the high computational cost and difficulty in balancing convergence efficiency of thresholding-based Newton-type methods for sparse optimization, this paper proposes a Compressed Newton Thresholding Optimization framework. It compresses the Newton direction onto a low-dimensional subspace and incorporates diagonal regularization, yielding two novel algorithms: Compressed Newton Hard Thresholding Pursuit (CNHTP) and Compressed Newton Optimal Thresholding Pursuit (CNOTP). Under the Restricted Isometry Property (RIP), we establish rigorous theoretical guarantees of global convergence. The algorithms integrate compressed direction computation, adaptive thresholding selection, and regularization, preserving Newton-type convergence rates while significantly reducing per-iteration complexity. Experiments demonstrate that the proposed methods match state-of-the-art algorithms in recovery success rate and solution accuracy, while achieving superior computational efficiency and robustness.
This work addresses the sparse solution recovery problem for linear systems involving concatenated orthogonal matrices. To exploit their structural properties, we propose a splitting-based alternating optimization framework—supporting both two-block and multi-block decompositions—that relies solely on matrix-vector products and low-dimensional orthogonal projections, thereby avoiding explicit matrix inversion or storage. By decomposing the large-scale system into coupled subsystems and solving them cooperatively via iterative updates, the algorithm is proven to converge globally to the sparse solution under a coherence constraint. Compared to mainstream iterative methods—including Orthogonal Matching Pursuit (OMP) and Iterative Shrinkage-Thresholding Algorithm (ISTA)—the proposed approach significantly reduces iteration counts while achieving faster convergence and enhanced numerical stability. This yields an efficient, scalable, and structurally aware paradigm for high-dimensional sparse signal recovery.