科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Tsinghua Science & Technology2025-12-19· Inference

Efficient Inference for Edge Large Language Models: A Survey

Guanyu Cai, Ruiming Tian, Lang Yang, Yunzhe Jia, Lingkun Li, Liang Wang

原始摘要(英文原文)· Original abstract
Abstract Large language models (LLMs) have demonstrated remarkable capabilities in natural language processing. Their massive computational and memory requirements often necessitate cloud-based deployment, introducing challenges related to cost, latency, privacy, and network reliability. Deploying on-device LLMs alleviates these challenges, but is hindered by the severe resource constraints of edge hardware. This survey reviews efficient inference techniques for edge LLMs, with a focus on two key strategies of speculative decoding and model offloading. We categorize strategies into single-device and multi-device types, systematically analyzing the principles, recent advancements, implementations, and support within edge frameworks. Finally, we highlight the open challenges and future research directions that will advance the field of edge LLM inference.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Efficient Inference for Edge Large Language Models: A Survey — 科研速览 Science Skim