Ke Lv, Sai Huang, Yuanyuan Yao, Weiwei Jiang, Zhiyong Feng
The artificial general intelligence (AGI) based on large language models (LLMs) and deep reinforcement learning (DRL) possesses cross-domain empowerment capabilities, offering task scheduling and resource allocation in mobile edge computing (MEC), presenting significant potential. This paper proposes an LLM-driven DRL framework named L2D2, which autonomously generates reward functions for DRL through LLMs and enables dynamic model optimization through a self-refinement loop mechanism. The reward function within the L2D2 framework can adjust the strategy based on environmental feedback, eliminating the need for manual redesign in complex low-altitude scenarios and reducing debugging costs. To validate the performance of L2D2, the framework is utilized in a multi-unmanned aerial vehicle (UAV)-assisted MEC heterogeneous network operating in the low-altitude airspace to enhance system energy efficiency. A novel dueling double deep Q-network (D3QN) is utilized as the DRL method within L2D2, named the L2D2-D3QN algorithm. To evaluate its effectiveness in enhancing system energy efficiency, a comprehensive comparison is conducted across various LLMs, including Deepseek-R1, GPT-4o, Llama-3.1-70B, Claude-3.7-Sonnet, and Qwen-2.5. The simulation results demonstrate that the L2D2-D3QN algorithm achieves up to 56% higher energy efficiency compared to DRL with human-designed reward function. Furthermore, the influence of LLM tokenization strategies on the performance of LLM-driven DRL is also explored.