科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Electronics2026-06-03· Benchmarking

Benchmarking Large Language Model Inference on Limited-Resource Edge Systems

Henrikas Giedra, Dalius Matuzevičius, Tomyslav Sledević, Giga Shubitidze, Artūras Serackis

原始摘要(英文原文)· Original abstract
Large language models (LLMs) are increasingly considered for deployment on edge and limited-resource systems, where local inference can reduce latency, improve privacy, and decrease dependence on cloud infrastructure. While prior studies have evaluated either task accuracy or hardware efficiency in isolation, few benchmarks combine generation-based response-quality evaluation with real-device power measurements on a representative limited-resource platform. This study addresses that gap by benchmarking twelve compact and mid-scale open-weight LLMs (sub-1B to 8B parameters), evaluating generation-based accuracy on a desktop platform and measuring deployment efficiency—throughput, power consumption, and energy use—on an NVIDIA Jetson Orin Nano Super; the accuracy–efficiency trade-off is thus established at the model-configuration level. Unlike prior Jetson-based evaluations relying solely on internal telemetry, this work pairs generation-compatible lm-eval accuracy tasks with a dual power-measurement setup that combines internal tegrastats rail readings with external board-level input power measured using a digital multimeter and explicitly compares GPU-accelerated and CPU-only inference modes. GPU-accelerated inference provided a clear advantage, increasing median throughput from 7.12 to 18.13 tok/s and improving external-meter energy efficiency from 0.453 to 0.823 tok/J, despite higher mean input power. Sub-1B models offered the best throughput and energy efficiency, whereas 7–8B models achieved stronger accuracy at a substantially higher energy cost per generated token. These results demonstrate that edge LLM deployment requires multi-objective evaluation balancing accuracy, throughput, power consumption, and energy efficiency.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Benchmarking Large Language Model Inference on Limited-Resource Edge Systems — 科研速览 Science Skim