Institution profile

NXAI

Industry researcheurope · at
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

KVpop -- Key-Value Cache Compression with Predictive Online Pruning

Jul 06, 2026

This work addresses the substantial memory and bandwidth bottlenecks in autoregressive decoding caused by the linear growth of key-value (KV) cache with context length. Existing KV cache eviction methods rely on static heuristics or proxy scores that inadequately estimate each cache entry’s contribution to future token generation, often resulting in significant performance degradation. To overcome this limitation, the authors propose a supervised learning framework that directly optimizes KV cache eviction under a fixed budget by leveraging future attention targets as supervision signals. They further introduce a delayed memory scorer that implicitly guides online pruning using near-future context, eliminating the need for explicit computation of dense attention maps. Evaluated on Qwen3-4B and Qwen3-8B, the method retains 97%–98% of original model performance at aggressive compression rates of 75%–88%, substantially outperforming current baselines.

0 citationsRead paper

Effective Distillation to Hybrid xLSTM Architectures

Mar 16, 2026

Existing methods struggle to effectively distill large language models based on quadratic attention into sub-quadratic linear architectures—such as xLSTM—without degrading downstream performance. This work proposes an efficient distillation framework tailored for xLSTM, which first employs a multi-expert linearized distillation phase followed by an expert-merging stage to consolidate multiple student experts into a single high-performance model. We introduce a lossless distillation evaluation metric based on tolerance-corrected win rates and tie rates, and achieve, for the first time, high-fidelity knowledge transfer from teacher models—including Llama, Qwen, and Olmo—to xLSTM. Experiments demonstrate that the proposed method not only recovers but often surpasses the teachers’ performance across diverse downstream tasks, significantly advancing the development of efficient alternatives to the Transformer architecture.

0 citationsRead paper

Rethinking Losses for Diffusion Bridge Samplers

Jun 12, 2025

This paper addresses the theoretical deficiency of the widely adopted Log Variance (LV) loss in diffusion bridge samplers—namely, its lack of rigorous derivation and misalignment between optimization objective and sampling performance. We propose rKL-LD, a novel loss function grounded in the reverse KL divergence and leveraging the log-derivative trick, which combines theoretical rigor with practical superiority. We first demonstrate that LV cannot be justified via the data processing inequality within the diffusion bridge framework, whereas rKL-LD emerges naturally and admits a clear variational interpretation. Experiments across multiple challenging benchmarks show that rKL-LD significantly improves sample quality (FID reduced by 12–28%), enhances training stability, and exhibits greater robustness to hyperparameter choices. Our work establishes a new theoretical standard and practical paradigm for diffusion bridge modeling.

0 citationsRead paper
Recent publications

Latest Papers

KVpop -- Key-Value Cache Compression with Predictive Online Pruning

Jul 06, 2026

This work addresses the substantial memory and bandwidth bottlenecks in autoregressive decoding caused by the linear growth of key-value (KV) cache with context length. Existing KV cache eviction methods rely on static heuristics or proxy scores that inadequately estimate each cache entry’s contribution to future token generation, often resulting in significant performance degradation. To overcome this limitation, the authors propose a supervised learning framework that directly optimizes KV cache eviction under a fixed budget by leveraging future attention targets as supervision signals. They further introduce a delayed memory scorer that implicitly guides online pruning using near-future context, eliminating the need for explicit computation of dense attention maps. Evaluated on Qwen3-4B and Qwen3-8B, the method retains 97%–98% of original model performance at aggressive compression rates of 75%–88%, substantially outperforming current baselines.

0 citationsRead paper

Effective Distillation to Hybrid xLSTM Architectures

Mar 16, 2026

Existing methods struggle to effectively distill large language models based on quadratic attention into sub-quadratic linear architectures—such as xLSTM—without degrading downstream performance. This work proposes an efficient distillation framework tailored for xLSTM, which first employs a multi-expert linearized distillation phase followed by an expert-merging stage to consolidate multiple student experts into a single high-performance model. We introduce a lossless distillation evaluation metric based on tolerance-corrected win rates and tie rates, and achieve, for the first time, high-fidelity knowledge transfer from teacher models—including Llama, Qwen, and Olmo—to xLSTM. Experiments demonstrate that the proposed method not only recovers but often surpasses the teachers’ performance across diverse downstream tasks, significantly advancing the development of efficient alternatives to the Transformer architecture.

0 citationsRead paper

Rethinking Losses for Diffusion Bridge Samplers

Jun 12, 2025

This paper addresses the theoretical deficiency of the widely adopted Log Variance (LV) loss in diffusion bridge samplers—namely, its lack of rigorous derivation and misalignment between optimization objective and sampling performance. We propose rKL-LD, a novel loss function grounded in the reverse KL divergence and leveraging the log-derivative trick, which combines theoretical rigor with practical superiority. We first demonstrate that LV cannot be justified via the data processing inequality within the diffusion bridge framework, whereas rKL-LD emerges naturally and admits a clear variational interpretation. Experiments across multiple challenging benchmarks show that rKL-LD significantly improves sample quality (FID reduced by 12–28%), enhances training stability, and exhibits greater robustness to hyperparameter choices. Our work establishes a new theoretical standard and practical paradigm for diffusion bridge modeling.

0 citationsRead paper