Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs"Difficult"Downstream Tasks in LLMs
This work challenges the prevailing assumption that small-magnitude weights in large language models (LLMs) are redundant, proposing instead the “Junk DNA Hypothesis”: low-magnitude weights encode essential knowledge for solving difficult downstream tasks. Method: We conduct systematic magnitude-based pruning—both structured and unstructured—alongside multi-granularity task difficulty quantification (e.g., reasoning depth, distribution shift, generalization gap), validated across model scales (7B–70B) and diverse benchmarks (MMLU, GSM8K, HumanEval). Contribution/Results: Pruning induces irreversible, monotonic performance degradation strictly correlated with task difficulty—degradation persists even after extensive fine-tuning—whereas quantization exhibits no such effect. This is the first study to empirically establish the functional necessity of small-magnitude weights from a task-difficulty perspective. We further propose novel, quantifiable cross-task difficulty metrics and demonstrate a strong negative correlation between optimal pruning ratio and task difficulty.