Real-Time dApps for AI-RAN: Measured Interface Requirements for Inline PHY and Slot-Level Control
本文研究了dApps在5G AI-RAN中的应用,通过测量不同框架的性能要求,提出了适用于实时物理层处理和时隙级控制的接口设计。
本文研究了dApps在5G AI-RAN中的应用,通过测量不同框架的性能要求,提出了适用于实时物理层处理和时隙级控制的接口设计。
本文介绍OCUDU dApp平台,通过开放运行时和E3接口,在5G分布式单元内实现实时AI-RAN应用的执行,解决缺乏开放式软件运行平台的问题。
本文针对6G中双侧AI模型的部署问题,通过将其集成到5G NR协议栈、优化模型以适应当前信道条件及采用零阶微调方法来减少训练和存储开销,提出了实用解决方案。
This work addresses the high latency and inefficiency in 5G physical layer and O-RAN fronthaul processing by proposing the first unified acceleration architecture that supports GPU-resident data. Built on CUDA, the OCUDU backend decouples acceleration interfaces to efficiently handle PDSCH, PUSCH, PRACH, SRS, split-8 lower-PHY transforms, and IQ compression/decompression. The architecture is compatible with both standard baseband processing and AI-RAN research, enabling concurrent execution of emerging algorithms such as neural receivers and AI-based channel estimation. Performance is optimized through a resource grid, device-side soft-bit buffers, CUDA streams and events, fixed scratchpads, and managed memory strategies, combined with slot-level batching and zero-copy mappings. Evaluated on an NVIDIA DGX Spark system, it achieves speedups of 91.4× for O-FH decompression, 28.8× for PRACH detection, 10.3× for PUSCH, and 2.7× for PDSCH over CPU baselines, with BLER performance degradation below 0.064 dB.
本文研究了dApps在5G AI-RAN中的应用,通过测量不同框架的性能要求,提出了适用于实时物理层处理和时隙级控制的接口设计。
本文介绍OCUDU dApp平台,通过开放运行时和E3接口,在5G分布式单元内实现实时AI-RAN应用的执行,解决缺乏开放式软件运行平台的问题。
本文针对6G中双侧AI模型的部署问题,通过将其集成到5G NR协议栈、优化模型以适应当前信道条件及采用零阶微调方法来减少训练和存储开销,提出了实用解决方案。
This work addresses the high latency and inefficiency in 5G physical layer and O-RAN fronthaul processing by proposing the first unified acceleration architecture that supports GPU-resident data. Built on CUDA, the OCUDU backend decouples acceleration interfaces to efficiently handle PDSCH, PUSCH, PRACH, SRS, split-8 lower-PHY transforms, and IQ compression/decompression. The architecture is compatible with both standard baseband processing and AI-RAN research, enabling concurrent execution of emerging algorithms such as neural receivers and AI-based channel estimation. Performance is optimized through a resource grid, device-side soft-bit buffers, CUDA streams and events, fixed scratchpads, and managed memory strategies, combined with slot-level batching and zero-copy mappings. Evaluated on an NVIDIA DGX Spark system, it achieves speedups of 91.4× for O-FH decompression, 28.8× for PRACH detection, 10.3× for PUSCH, and 2.7× for PDSCH over CPU baselines, with BLER performance degradation below 0.064 dB.