Institution profile

DeepSig, Inc.

Industry researchnorthamerica · us
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

GPU-Resident CUDA Acceleration for OCUDU 5G PHY and O-RAN Fronthaul: Architecture and Preliminary Performance

Aug 04, 2026

This work addresses the high latency and inefficiency in 5G physical layer and O-RAN fronthaul processing by proposing the first unified acceleration architecture that supports GPU-resident data. Built on CUDA, the OCUDU backend decouples acceleration interfaces to efficiently handle PDSCH, PUSCH, PRACH, SRS, split-8 lower-PHY transforms, and IQ compression/decompression. The architecture is compatible with both standard baseband processing and AI-RAN research, enabling concurrent execution of emerging algorithms such as neural receivers and AI-based channel estimation. Performance is optimized through a resource grid, device-side soft-bit buffers, CUDA streams and events, fixed scratchpads, and managed memory strategies, combined with slot-level batching and zero-copy mappings. Evaluated on an NVIDIA DGX Spark system, it achieves speedups of 91.4× for O-FH decompression, 28.8× for PRACH detection, 10.3× for PUSCH, and 2.7× for PDSCH over CPU baselines, with BLER performance degradation below 0.064 dB.

0 citationsRead paper
Recent publications

Latest Papers

GPU-Resident CUDA Acceleration for OCUDU 5G PHY and O-RAN Fronthaul: Architecture and Preliminary Performance

Aug 04, 2026

This work addresses the high latency and inefficiency in 5G physical layer and O-RAN fronthaul processing by proposing the first unified acceleration architecture that supports GPU-resident data. Built on CUDA, the OCUDU backend decouples acceleration interfaces to efficiently handle PDSCH, PUSCH, PRACH, SRS, split-8 lower-PHY transforms, and IQ compression/decompression. The architecture is compatible with both standard baseband processing and AI-RAN research, enabling concurrent execution of emerging algorithms such as neural receivers and AI-based channel estimation. Performance is optimized through a resource grid, device-side soft-bit buffers, CUDA streams and events, fixed scratchpads, and managed memory strategies, combined with slot-level batching and zero-copy mappings. Evaluated on an NVIDIA DGX Spark system, it achieves speedups of 91.4× for O-FH decompression, 28.8× for PRACH detection, 10.3× for PUSCH, and 2.7× for PDSCH over CPU baselines, with BLER performance degradation below 0.064 dB.

0 citationsRead paper