π€ AI Summary
This work proposes the first all-optical convolutional neural network that operates entirely in the optical domain without requiring optoelectronic conversion, addressing the throughput and power bottlenecks of conventional electronic CNNs in image processing. Implemented on a silicon photonic platform, the system integrates MachβZehnder interferometer (MZI) meshes for linear operations, wavelength-division multiplexing for pooling, and microring resonators as nonlinear activation units. A hybrid training strategy combines a differentiable numerical twin model with an in-situ simultaneous perturbation stochastic approximation (SPSA) optimization algorithm. Evaluated on MNIST image classification, the system achieves 94% accuracy, exhibits remarkable robustness with only a 0.43% performance drop under severe thermal crosstalk, and demonstrates 100β242Γ higher energy efficiency per inference compared to state-of-the-art electronic GPUs.
π Abstract
Photonic computing is a computing paradigm which have great potential to overcome the energy bottlenecks of electronic von Neumann architecture. Throughput and power consumption are fundamental limitations of Complementary-metal-oxide-semiconductor (CMOS) chips, therefore convolutional neural network (CNN) is revolutionising machine learning, computer vision and other image based applications. In this work, we propose and validate a fully photonic convolutional neural network (PCNN) that performs MNIST image classification entirely in the optical domain, achieving 94 percent test accuracy. Unlike existing architectures that rely on frequent in-between conversions from optical to electrical and back to optical (O/E/O), our system maintains coherent processing utilizing Mach-Zehnder interferometer (MZI) meshes, wavelength-division multiplexed (WDM) pooling, and microring resonator-based nonlinearities. The max pooling unit is fully implemented on silicon photonics, which does not require opto-electrical or electrical conversions. To overcome the challenges of training physical phase shifter parameters, we introduce a hybrid training methodology deploying a mathematically exact differentiable digital twin for ex-situ backpropagation, followed by in-situ fine-tuning via Simultaneous Perturbation Stochastic Approximation (SPSA) algorithm. Our evaluation demonstrates significant robustness to thermal crosstalk (only 0.43 percent accuracy degradation at severe coupling) and achieves 100 to 242 times better energy efficiency than state-of-the-art electronic GPUs for single-image inference.