Institution profile

Rain AI

Industry researchnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Training a Predictive Coding Network on ImageNet using Equilibrium Propagation

Jun 02, 2026

This work addresses the limited scalability of Predictive Coding Networks (PCNs) and Equilibrium Propagation (EP) in large-scale vision tasks by introducing a novel hybrid approach that integrates key innovations from both frameworks. Specifically, it proposes an energy-based PCN architecture, a new equilibrium mechanism tailored for PCNs, and a centralized variant of EP. This combination enables, for the first time, the successful training of a 10-layer convolutional PCN (VGG10) on ImageNet. The method substantially improves scalability, achieving a Top-5 error rate of 13.23%—closely approaching the performance of standard backpropagation baselines (12.2%)—thereby demonstrating its effectiveness and practical viability for large-scale visual recognition.

0 citationsRead paper

EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models

Apr 13, 2026

This work addresses the memory bandwidth bottleneck that limits the energy efficiency and throughput of autoregressive inference with small language models (SLMs) on edge devices. To overcome this challenge, the authors propose EdgeCIM, the first framework integrating compute-in-memory (CIM) with holistic hardware-software co-design. By leveraging a 65nm CIM macro, INT4 quantization, and a tiling-aware weight mapping strategy, EdgeCIM optimizes the end-to-end decoding pipeline to alleviate DRAM bottlenecks. Evaluated on LLaMA3.2-1B, EdgeCIM achieves a 7.3× higher throughput—reaching 336.42 tokens/s—and a 49.59× improvement in energy efficiency, attaining 173.02 tokens/J, compared to NVIDIA Orin Nano. The design also enables exploration of models up to 4 billion parameters within the edge deployment regime.

0 citationsRead paper

Mixed-Precision Quantization for Language Models: Techniques and Prospects

Oct 19, 2025

Scaling large language models (LLMs) leads to unsustainable computational, memory, and energy costs. This paper proposes a hardware-aware mixed-precision quantization system for LLMs: it establishes a unified taxonomy and systematically analyzes bit-width allocation across weights, activations, and KV caches; designs hardware-efficient uniform and non-uniform quantizers integrating post-training quantization, fine-grained quantization control, and scalable optimization algorithms to support sub-INT8 compression. Compared to uniform quantization, our approach reduces memory footprint by up to 62% and inference latency by up to 3.1×, while preserving perplexity and zero-shot task accuracy. The core contribution lies in empirically characterizing the heterogeneous precision sensitivity of LLM components and delivering the first holistic mixed-precision quantization framework that jointly addresses theoretical analysis, systems implementation, and hardware co-design.

0 citationsRead paper
Recent publications

Latest Papers

Training a Predictive Coding Network on ImageNet using Equilibrium Propagation

Jun 02, 2026

This work addresses the limited scalability of Predictive Coding Networks (PCNs) and Equilibrium Propagation (EP) in large-scale vision tasks by introducing a novel hybrid approach that integrates key innovations from both frameworks. Specifically, it proposes an energy-based PCN architecture, a new equilibrium mechanism tailored for PCNs, and a centralized variant of EP. This combination enables, for the first time, the successful training of a 10-layer convolutional PCN (VGG10) on ImageNet. The method substantially improves scalability, achieving a Top-5 error rate of 13.23%—closely approaching the performance of standard backpropagation baselines (12.2%)—thereby demonstrating its effectiveness and practical viability for large-scale visual recognition.

0 citationsRead paper

EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models

Apr 13, 2026

This work addresses the memory bandwidth bottleneck that limits the energy efficiency and throughput of autoregressive inference with small language models (SLMs) on edge devices. To overcome this challenge, the authors propose EdgeCIM, the first framework integrating compute-in-memory (CIM) with holistic hardware-software co-design. By leveraging a 65nm CIM macro, INT4 quantization, and a tiling-aware weight mapping strategy, EdgeCIM optimizes the end-to-end decoding pipeline to alleviate DRAM bottlenecks. Evaluated on LLaMA3.2-1B, EdgeCIM achieves a 7.3× higher throughput—reaching 336.42 tokens/s—and a 49.59× improvement in energy efficiency, attaining 173.02 tokens/J, compared to NVIDIA Orin Nano. The design also enables exploration of models up to 4 billion parameters within the edge deployment regime.

0 citationsRead paper

Mixed-Precision Quantization for Language Models: Techniques and Prospects

Oct 19, 2025

Scaling large language models (LLMs) leads to unsustainable computational, memory, and energy costs. This paper proposes a hardware-aware mixed-precision quantization system for LLMs: it establishes a unified taxonomy and systematically analyzes bit-width allocation across weights, activations, and KV caches; designs hardware-efficient uniform and non-uniform quantizers integrating post-training quantization, fine-grained quantization control, and scalable optimization algorithms to support sub-INT8 compression. Compared to uniform quantization, our approach reduces memory footprint by up to 62% and inference latency by up to 3.1×, while preserving perplexity and zero-shot task accuracy. The core contribution lies in empirically characterizing the heterogeneous precision sensitivity of LLM components and delivering the first holistic mixed-precision quantization framework that jointly addresses theoretical analysis, systems implementation, and hardware co-design.

0 citationsRead paper