Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention
This work addresses the challenges of transferring adaptive feature retention (AFR) from unstructured to structured pruning, which include heterogeneous pruning score distributions, loss of sign information, and outlier interference. To bridge this gap, the authors propose a unified structured pruning framework that introduces power transformation to align score distributions, designs a sign-preserving aggregation mechanism to maintain consistent optimization directions, and incorporates a percentile-based outlier removal strategy. This approach effectively narrows the performance gap between structured and unstructured pruning, achieving accuracy on Llama-3-8B, Vicuna-v1.5-13B, and LLaVA-v1.5-13B models that closely matches unstructured pruning while delivering substantial real-world inference speedups.