Decoder-Agnostic Token Merging for Vision Transformers: A Systematic Study of G2TM
研究通过Graph-Guided Token Merging(G2TM)方法减少Vision Transformers的计算成本,证明其有效性主要取决于编码器而非解码器。
研究通过Graph-Guided Token Merging(G2TM)方法减少Vision Transformers的计算成本,证明其有效性主要取决于编码器而非解码器。
为解决驾驶图像照明修改难题,HELIOS提出一种基于无标签真实数据的新方法,通过整合反照率条件和循环一致扩散管道等技术,实现日夜连续场景照明调整。
This work addresses the problem of efficient online change-point detection in both univariate and multivariate data streams by introducing a novel method grounded in the Focus algorithm family. Leveraging the generalized likelihood ratio test, the approach enables exact detection of a single change point without requiring approximations. By exploiting the relationship between candidate change-point locations and the geometric structure of the data, it achieves a computational complexity of approximately $\log(n)^d$ per iteration. Notably, this is the first method to support exponential-family models, nonparametric settings, and autoregressive data under no approximation assumptions, integrating natural exponential-family modeling, empirical cumulative distribution functions, and geometric optimization techniques. The accompanying R/Python software package substantially enhances the efficiency and applicability of change-point detection in high-dimensional streaming data.
This work addresses the challenge of achieving conditional coverage in conformal prediction without relying on strong structural assumptions. To this end, it proposes the PIT-CP method, which post-processes arbitrary nonconformity scores via one-dimensional conditional density estimation—using tools such as mixture density networks or conditional normalizing flows—to map them into approximately feature-independent pivotal scores. The approach requires no stringent modeling assumptions while preserving marginal coverage and the geometric structure of prediction sets, and it substantially improves conditional coverage performance. Theoretical analysis provides both deterministic and high-probability upper bounds on the conditional coverage gap, along with formal guarantees on the volume and symmetric difference of the resulting prediction sets.
This work addresses the high computational and memory costs incurred by vision-language models when processing long visual token sequences. Existing pruning methods, which rely on local heuristics, often suffer from positional bias and fragmented information retention, making it difficult to preserve critical semantics under high compression ratios. To overcome these limitations, the authors propose a training-free, plug-and-play pruning approach that introduces singular value decomposition (SVD) into visual token selection for the first time. By leveraging statistical leverage scores, the method identifies the top-K tokens that contribute most significantly to the global principal components. This strategy effectively circumvents the shortcomings of local heuristics and achieves substantial performance gains over existing techniques—even under extreme compression settings retaining only 16 or 32 tokens—while maintaining strong model performance on detail-rich images.
研究通过Graph-Guided Token Merging(G2TM)方法减少Vision Transformers的计算成本,证明其有效性主要取决于编码器而非解码器。
为解决驾驶图像照明修改难题,HELIOS提出一种基于无标签真实数据的新方法,通过整合反照率条件和循环一致扩散管道等技术,实现日夜连续场景照明调整。
This work addresses the problem of efficient online change-point detection in both univariate and multivariate data streams by introducing a novel method grounded in the Focus algorithm family. Leveraging the generalized likelihood ratio test, the approach enables exact detection of a single change point without requiring approximations. By exploiting the relationship between candidate change-point locations and the geometric structure of the data, it achieves a computational complexity of approximately $\log(n)^d$ per iteration. Notably, this is the first method to support exponential-family models, nonparametric settings, and autoregressive data under no approximation assumptions, integrating natural exponential-family modeling, empirical cumulative distribution functions, and geometric optimization techniques. The accompanying R/Python software package substantially enhances the efficiency and applicability of change-point detection in high-dimensional streaming data.
This work addresses the challenge of achieving conditional coverage in conformal prediction without relying on strong structural assumptions. To this end, it proposes the PIT-CP method, which post-processes arbitrary nonconformity scores via one-dimensional conditional density estimation—using tools such as mixture density networks or conditional normalizing flows—to map them into approximately feature-independent pivotal scores. The approach requires no stringent modeling assumptions while preserving marginal coverage and the geometric structure of prediction sets, and it substantially improves conditional coverage performance. Theoretical analysis provides both deterministic and high-probability upper bounds on the conditional coverage gap, along with formal guarantees on the volume and symmetric difference of the resulting prediction sets.
This work addresses the high computational and memory costs incurred by vision-language models when processing long visual token sequences. Existing pruning methods, which rely on local heuristics, often suffer from positional bias and fragmented information retention, making it difficult to preserve critical semantics under high compression ratios. To overcome these limitations, the authors propose a training-free, plug-and-play pruning approach that introduces singular value decomposition (SVD) into visual token selection for the first time. By leveraging statistical leverage scores, the method identifies the top-K tokens that contribute most significantly to the global principal components. This strategy effectively circumvents the shortcomings of local heuristics and achieves substantial performance gains over existing techniques—even under extreme compression settings retaining only 16 or 32 tokens—while maintaining strong model performance on detail-rich images.