FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大型语言模型在资源受限设备上的部署问题,提出了一种基于Fisher信息的自适应混合精度权重量化方法FAMPWQ,通过层自适应量化有效提升了推理性能。
📝 Abstract
Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of LLMs hinder the deployment on resource-constrained devices. Although model quantization stands out as an effective approach, conventional quantization approaches typically incur severe performance degradation due to uniform bit-width or simple heuristic sensitivity evaluation. In this paper, we propose a novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs. First, we propose a system model with a novel Fisher information metric to measure the layer-wise sensitivity to quantization. Second, we propose a reinforcement learning-based bit-width allocator in FAMPWQ, which generates an adaptive bit-width allocation strategy based on the Fisher information sensitivity metric. Extensive experiments on 7 models and 5 benchmarks demonstrate that FAMPWQ significantly outperforms 7 baseline approaches in terms of PPL (up to 3.39 smaller), accuracy (up to 6.87% higher), and LLM-as-a-judge comparison (up to 76% win rate).
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
resource requirements
model quantization
performance degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fisher Information
Adaptive Mixed Precision
Weight Quantization
Reinforcement Learning
Layer-wise Sensitivity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.