FairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出FairCompressAgent框架,通过集成多种压缩技术并结合语言模型规划器来解决公平性感知模型压缩的问题,以平衡准确性、公平性和部署成本。
📝 Abstract
Fairness-aware model compression requires selecting methods and configurations that balance accuracy, fairness, and deployment cost. These decisions become more difficult when compression methods are composed or the user's requirements change. In this paper, we propose FairCompressAgent (FCA), an agentic framework that integrates fairness-aware pruning, incremental quantization, and sparse low-rank factorization through a common operator interface. A language-model planner uses model profiles and measured outcomes to select compression configurations, while an execution layer performs compression, fine-tuning, evaluation, and constraint-based selection. FCA also supports requirement updates and reports the remaining violation when a request cannot be satisfied. Experiments on Fitzpatrick-17k with VGG-11 compare four search methods over 40 measured configurations. Under the accuracy-constrained request, FCA selects a compressed model with 59.54% less inference tensor storage, while validation average precision increases from 0.5141 to 0.5233 and equalized opportunity (EOpp) decreases from 0.2251 to 0.2168. It reaches the same final selection as one-shot planning with 7.33 versus 12 candidate evaluations on average, under their respective stopping policies. Repeated fine-tuning, held-out testing, and online requirement updates characterize the stability and interactive use of this compression workflow. The results demonstrate how measured feedback and explicit constraints support the selection and interactive refinement of fairness-aware compression configurations.
Problem

Research questions and friction points this paper is trying to address.

fairness-aware model compression
deployment cost
user's requirements
Innovation

Methods, ideas, or system contributions that make the work stand out.

fairness-aware model compression
incremental quantization
sparse low-rank factorization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuanbo Guo
Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, USA
Yiyu Shi
Yiyu Shi
Full Professor, University of Notre Dame
hardware/software co-designdeep learning accelerationon-device AIAI for healthcare