Institution profile

RIKEN Center for Computational Science

Academic institutionasia · jp
Official website
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

ZIPBrain: Can EEG Foundation Models Be Faster, Locally Deployable, but Accurate?

Aug 07, 2026

This work addresses the quadratic computational overhead and redundancy in EEG foundation models caused by long input sequences and low signal-to-noise ratios. To this end, the authors propose ZIPBrain, a training-free, plug-and-play perceptual redundancy-aware token pooling module. ZIPBrain partitions EEG tokens into redundant and distinctive groups, then merges each redundant token with its most similar distinctive counterpart, substantially compressing sequence length. The method seamlessly integrates into standard Transformer encoders without requiring fine-tuning and leverages CUDA Graph acceleration for efficient inference. Evaluated across multiple EEG foundation models, ZIPBrain consistently improves average accuracy by 1.3%–10.5% while reducing inference time by 32.7% on average—up to 41.8%—achieving a favorable trade-off between efficiency and performance.

0 citationsRead paper

Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers

Aug 07, 2026

This work addresses the challenges of scaling flexible macromolecular docking on GPU-based supercomputers, which are hindered by irregular computation, low parallelism, and load imbalance. The authors propose SparkleDock, a novel framework that restructures the Glowworm Swarm Optimization (GSO) algorithm to expose fine-grained parallelism at the individual level and reformulates energy scoring as matrix operations amenable to Tensor Cores. Coupled with a performance-model-driven scheduling strategy, SparkleDock achieves cross-GPU load balancing and out-of-core scalability. For the first time, this approach enables near-real-time GSO-based flexible docking on GPU supercomputers: achieving 9.7× and 18.9× speedups on a single A100 and H100 GPU, respectively, and reducing docking time from hours to seconds at a scale of 512 GPUs—delivering over two orders of magnitude overall acceleration.

0 citationsRead paper
Recent publications

Latest Papers

ZIPBrain: Can EEG Foundation Models Be Faster, Locally Deployable, but Accurate?

Aug 07, 2026

This work addresses the quadratic computational overhead and redundancy in EEG foundation models caused by long input sequences and low signal-to-noise ratios. To this end, the authors propose ZIPBrain, a training-free, plug-and-play perceptual redundancy-aware token pooling module. ZIPBrain partitions EEG tokens into redundant and distinctive groups, then merges each redundant token with its most similar distinctive counterpart, substantially compressing sequence length. The method seamlessly integrates into standard Transformer encoders without requiring fine-tuning and leverages CUDA Graph acceleration for efficient inference. Evaluated across multiple EEG foundation models, ZIPBrain consistently improves average accuracy by 1.3%–10.5% while reducing inference time by 32.7% on average—up to 41.8%—achieving a favorable trade-off between efficiency and performance.

0 citationsRead paper

Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers

Aug 07, 2026

This work addresses the challenges of scaling flexible macromolecular docking on GPU-based supercomputers, which are hindered by irregular computation, low parallelism, and load imbalance. The authors propose SparkleDock, a novel framework that restructures the Glowworm Swarm Optimization (GSO) algorithm to expose fine-grained parallelism at the individual level and reformulates energy scoring as matrix operations amenable to Tensor Cores. Coupled with a performance-model-driven scheduling strategy, SparkleDock achieves cross-GPU load balancing and out-of-core scalability. For the first time, this approach enables near-real-time GSO-based flexible docking on GPU supercomputers: achieving 9.7× and 18.9× speedups on a single A100 and H100 GPU, respectively, and reducing docking time from hours to seconds at a scale of 512 GPUs—delivering over two orders of magnitude overall acceleration.

0 citationsRead paper