科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ IEEE Internet of Things Journal2026-01-23· Computer science

Adaptive Pruning for Large Language Models With Structural Importance Awareness

Haotian Zheng, Jinke Ren, Yatong Han, Yushan Sun, Ruichen Zhang, Wenbo Zhang, Zhen Li, Dusit Niyato, Shuguang Cui

原始摘要(英文原文)· Original abstract
The recent advancements in large language models (LLMs) have significantly enhanced language understanding and content generation capabilities. However, the deployment of LLMs on resource-constrained Internet of Things (IoT) devices remains challenging due to their substantial computational and storage requirements. To address this issue, we propose a novel LLM pruning method, termed structurally-aware adaptive pruning (SAAP), to reduce computational and storage costs for LLMs while maintaining model performance. Specifically, SAAP first leverages maximum likelihood estimation to calibrate traditional structural importance metrics for LLM pruning. Next, it employs a Bayesian fusion approach to address the predictive uncertainty in multi-granularity metrics, enabling accurate assessments of structural importance for LLMs. Then, SAAP introduces a cross-layer importance alignment mechanism based on quantile mapping, which normalizes layer-wise importance scores to ensure consistent pruning from a global perspective. Furthermore, SAAP develops an efficient block-wise fine-tuning strategy for enhancing the performance of the LLM after pruning. To validate the effectiveness of SAAP, we conduct extensive experiments on nine open-source LLMs across two representative tasks—language modeling and zero-shot classification. Experimental results show that SAAP consistently outperforms several baseline methods, achieving accuracy improvements of 2.5%, 2.63%, and 2.44% on LLaMA-7B, Vicuna-7B, and LLaMA-13B when the pruning ratio is 50%. Finally, SAAP is implemented on a testbed—NVIDIA Jetson AGX Orin 32GB Developer Kit. Test results demonstrate that compared to the foundation LLM, SAAP enhances the inference speed by 86.86% at a pruning ratio of 50%, highlighting its potential for practical deployment on resource-constrained IoT devices.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Adaptive Pruning for Large Language Models With Structural Importance Awareness — 科研速览 Science Skim