🤖 AI Summary
Existing stochastic computers face two critical bottlenecks: limited probabilistic modeling capability and highly customized, non-scalable hardware architectures—severely hindering practical deployment. This paper introduces a fully transistor-based probabilistic computing architecture that, for the first time, directly maps denoising diffusion models—particularly those with strong multimodal modeling capacity—onto scalable, all-CMOS hardware platforms. Departing from conventional reliance on dedicated random-number generators and stochastic logic gates, the architecture enables probabilistic signal processing at the transistor level, synergistically co-optimized across the system stack to achieve high energy efficiency in hardware-accelerated model execution. Experimental evaluation on standard image benchmarks demonstrates performance parity with GPUs while reducing energy consumption by four orders of magnitude (≈10⁴×). This work fundamentally resolves the long-standing trade-off between modeling expressivity and hardware scalability in stochastic computing.
📝 Abstract
The proliferation of probabilistic AI has promoted proposals for specialized stochastic computers. Despite promising efficiency gains, these proposals have failed to gain traction because they rely on fundamentally limited modeling techniques and exotic, unscalable hardware. In this work, we address these shortcomings by proposing an all-transistor probabilistic computer that implements powerful denoising models at the hardware level. A system-level analysis indicates that devices based on our architecture could achieve performance parity with GPUs on a simple image benchmark using approximately 10,000 times less energy.