High-Rate Nested-Lattice Quantized Matrix Multiplication with Small Lookup Tables

📅 2025-05-19

📈 Citations: 0

✨ Influential: 0

career value

197K/year

🤖 AI Summary

To address the prohibitively large lookup table (LUT) overhead in high-rate nested lattice quantization—where LUT size scales exponentially with rate as $2^{2dR}$ and is tightly coupled to the quantization rate—this paper proposes a hierarchical nested lattice quantization framework. The $d$-dimensional vector quantization is decomposed into $M$ layers, each employing a nested lattice structure combined with product codes. This reduces the per-layer LUT size to $2^{2dR/M}$, effectively decoupling LUT complexity from the overall rate. To our knowledge, this is the first scheme to break the linear dependence between LUT size and quantization rate while preserving asymptotically negligible distortion penalty. Numerical experiments confirm that the method achieves near-optimal quantization accuracy at high rates, compressing the total LUT size by a factor of $M$-th root relative to conventional approaches. The resulting reduction significantly enhances hardware efficiency and deployment feasibility.

Technology Category

Application Category

📝 Abstract

Recent work have shown that the quantization for matrix multiplication problem can be optimally solved by quantizing each column in each matrix using a nested lattice code, and then multiplying the de-quantized matrices. It was further demonstrated that when product codes of sub-dimension $d$ and rate $R$ are used, the de-quantization and inner product operations can be implemented with querying a lookup table (LUT) of size $2^{2dR}$, but this is only useful when $dR$ is sufficiently small. This in turn limits LUT-based inner product decoding to low-rate quantizers. In this work, we develop a rate $R$ hierarchical nested lattice quantization framework, which quantizes each vector to $M$ layers, and admits LUT-based inner product decoding using an LUT of size $2^{2dfrac{R}{M}}$, allowing for high-rate quantization. We provide analytic bounds on the loss of the developed scheme compared to standard nested lattice quantizers, and also numerically illustrate that this loss is negligible. Thus, our scheme enables to use small LUTs without compromising the overall distortion.

Problem

Research questions and friction points this paper is trying to address.

Enables high-rate quantization with small lookup tables

Reduces LUT size for inner product decoding

Minimizes distortion in nested lattice quantization

Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical nested lattice quantization framework

Small lookup tables for high-rate quantization

Negligible distortion compared to standard quantizers

🔎 Similar Papers

Fast Matrix Multiplications for Lookup Table-Quantized LLMs

2024-07-15Conference on Empirical Methods in Natural Language ProcessingCitations: 8

Qualcomm

$140,800.00 - $211,200.00

San Diego, California, United States of America

Authors to Follow