What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过不同架构和任务复杂度评估剪枝对智能家居工具调用的影响,发现密集模型在剪枝后性能迅速下降,而MoE模型更耐剪枝。
📝 Abstract
Pruning can reduce the deployment cost of large language models (LLMs), but its impact on context-grounded tool calling remains poorly understood. We systematically study pruning-induced degradation in smart-home tool calling across four LLMs spanning dense Transformer, dense hybrid, and mixture-of-experts (MoE) architectures, together with depth, width, hybrid, and expert pruning methods. After post-pruning supervised fine-tuning (SFT), we evaluate more than 19,500 instances from three smart-home datasets. Beyond aggregate task accuracy, we characterize degradation along two dimensions: action components (i.e., operation, device, argument, and value) and task complexity. Our results show that dense models have narrow safe pruning regions followed by sharp degradation, while MoE models tolerate substantially more pruning. Pruning degrades grounded specificity before schema-level intent, and aggressive dense pruning can induce systematic over-refusal. These findings highlight the importance of evaluating pruning beyond aggregate accuracy when selecting pruned LLMs for reliable tool execution.
Problem

Research questions and friction points this paper is trying to address.

Pruning
Smart Home
Large Language Models
Task Complexity
Performance Degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

pruning
smart-home tool calling
model architectures
task complexity
degradation analysis
🔎 Similar Papers
No similar papers found.
C
Congjing Zhang
Alexa Home AI, Amazon.com; University of Washington
V
Vashishtha Patil
Alexa Home AI, Amazon.com
Henning Lange
Henning Lange
Alexa Home AI, Amazon.com
U
Usman Aleem
Alexa Home AI, Amazon.com