An Efficient Out-of-Core Tomographic Imaging Framework for Edge Devices
本文提出了一种名为edgeFBP的框架,通过采用混合精度策略加速反投影内核,解决了边缘设备上因计算能力和内存限制而难以进行CT成像的问题。
本文提出了一种名为edgeFBP的框架,通过采用混合精度策略加速反投影内核,解决了边缘设备上因计算能力和内存限制而难以进行CT成像的问题。
该研究通过使用实地实验数据微调大型语言模型来预测游客轨迹,解决了传统方法难以泛化到未观察场景的问题。
研究通过微调语言模型来模拟人群行为,使用迭代比例拟合方法调整目的地分布以匹配观察数据,解决了个体行为不确定问题。
This work addresses the quadratic computational overhead and redundancy in EEG foundation models caused by long input sequences and low signal-to-noise ratios. To this end, the authors propose ZIPBrain, a training-free, plug-and-play perceptual redundancy-aware token pooling module. ZIPBrain partitions EEG tokens into redundant and distinctive groups, then merges each redundant token with its most similar distinctive counterpart, substantially compressing sequence length. The method seamlessly integrates into standard Transformer encoders without requiring fine-tuning and leverages CUDA Graph acceleration for efficient inference. Evaluated across multiple EEG foundation models, ZIPBrain consistently improves average accuracy by 1.3%–10.5% while reducing inference time by 32.7% on average—up to 41.8%—achieving a favorable trade-off between efficiency and performance.
This work addresses the challenges of scaling flexible macromolecular docking on GPU-based supercomputers, which are hindered by irregular computation, low parallelism, and load imbalance. The authors propose SparkleDock, a novel framework that restructures the Glowworm Swarm Optimization (GSO) algorithm to expose fine-grained parallelism at the individual level and reformulates energy scoring as matrix operations amenable to Tensor Cores. Coupled with a performance-model-driven scheduling strategy, SparkleDock achieves cross-GPU load balancing and out-of-core scalability. For the first time, this approach enables near-real-time GSO-based flexible docking on GPU supercomputers: achieving 9.7× and 18.9× speedups on a single A100 and H100 GPU, respectively, and reducing docking time from hours to seconds at a scale of 512 GPUs—delivering over two orders of magnitude overall acceleration.
本文提出了一种名为edgeFBP的框架,通过采用混合精度策略加速反投影内核,解决了边缘设备上因计算能力和内存限制而难以进行CT成像的问题。
该研究通过使用实地实验数据微调大型语言模型来预测游客轨迹,解决了传统方法难以泛化到未观察场景的问题。
研究通过微调语言模型来模拟人群行为,使用迭代比例拟合方法调整目的地分布以匹配观察数据,解决了个体行为不确定问题。
This work addresses the quadratic computational overhead and redundancy in EEG foundation models caused by long input sequences and low signal-to-noise ratios. To this end, the authors propose ZIPBrain, a training-free, plug-and-play perceptual redundancy-aware token pooling module. ZIPBrain partitions EEG tokens into redundant and distinctive groups, then merges each redundant token with its most similar distinctive counterpart, substantially compressing sequence length. The method seamlessly integrates into standard Transformer encoders without requiring fine-tuning and leverages CUDA Graph acceleration for efficient inference. Evaluated across multiple EEG foundation models, ZIPBrain consistently improves average accuracy by 1.3%–10.5% while reducing inference time by 32.7% on average—up to 41.8%—achieving a favorable trade-off between efficiency and performance.
This work addresses the challenges of scaling flexible macromolecular docking on GPU-based supercomputers, which are hindered by irregular computation, low parallelism, and load imbalance. The authors propose SparkleDock, a novel framework that restructures the Glowworm Swarm Optimization (GSO) algorithm to expose fine-grained parallelism at the individual level and reformulates energy scoring as matrix operations amenable to Tensor Cores. Coupled with a performance-model-driven scheduling strategy, SparkleDock achieves cross-GPU load balancing and out-of-core scalability. For the first time, this approach enables near-real-time GSO-based flexible docking on GPU supercomputers: achieving 9.7× and 18.9× speedups on a single A100 and H100 GPU, respectively, and reducing docking time from hours to seconds at a scale of 512 GPUs—delivering over two orders of magnitude overall acceleration.