Institution profile

A*STAR Centre for Frontier AI Research

Academic institutionasia · sg
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

Aug 12, 2026

This work addresses the inefficiency of existing parallel decoding strategies in diffusion language models, which overlook the potential of early deterministic decisions to enhance global decoding efficiency. The authors propose a training-free active parallel decoding method that, for the first time, identifies and leverages a “ripple effect” during decoding: by detecting medium-entropy “pivot” positions, prospectively evaluating their impact on downstream uncertainty, and dynamically scheduling optimal decoding paths using KV cache management. Evaluated across three diffusion language models and four benchmarks spanning reasoning and code generation, the approach achieves 4–10× end-to-end speedup (up to 18× in peak cases) while preserving generation quality and consistently outperforming prior state-of-the-art baselines by up to 5.49% in accuracy across most settings.

0 citationsRead paper
Recent publications

Latest Papers

Ripple-Pivot Search: Active Parallel Decoding for Diffusion Large Language Models

Aug 12, 2026

This work addresses the inefficiency of existing parallel decoding strategies in diffusion language models, which overlook the potential of early deterministic decisions to enhance global decoding efficiency. The authors propose a training-free active parallel decoding method that, for the first time, identifies and leverages a “ripple effect” during decoding: by detecting medium-entropy “pivot” positions, prospectively evaluating their impact on downstream uncertainty, and dynamically scheduling optimal decoding paths using KV cache management. Evaluated across three diffusion language models and four benchmarks spanning reasoning and code generation, the approach achieves 4–10× end-to-end speedup (up to 18× in peak cases) while preserving generation quality and consistently outperforming prior state-of-the-art baselines by up to 5.49% in accuracy across most settings.

0 citationsRead paper