Yuzhu Liang, Changfu Xu, Yaxin Mei, Haodong Zou, J. Y. Guo, Xinggang Fan, Tian Wang, Haiyang Huang
Large Language Models (LLMs) have achieved remarkable performance across various artificial intelligence applications. However, current LLMs cannot be deployed directly on edge nodes due to their large number of parameters. Fortunately, model compression technology has been proposed to reduce the computational workload and memory usage of LLMs, enabling further edge-based LLM services. However, existing research typically concentrates on isolated compression algorithms and lacks a comprehensive perspective on how to leverage these techniques for practical, end-to-end LLM deployment in edge environments. In this survey, we review edge-oriented LLM compression techniques and software–hardware co-design strategies to enable efficient LLM deployment on resource-constrained edge systems and guide future research in this area. First, we analyze techniques for LLM compression from the perspective of cloud–edge collaborative intelligence, including model quantization, parameter pruning, and knowledge distillation. Second, we present several hybrid model frameworks tailored to dynamic, heterogeneous edge environments, based on model architecture, application scenarios, and combination selection. Third, we further refine a four-layer software–hardware codesign and an overhead-aware LLM deployment optimization. Finally, we discuss the challenges of current model compression approaches and offer insights into future research directions, with a focus on edge-based LLM services.