Yuqi Li, Zijie Zhou, Zhiyuan Peng, Junhao Dong, Haochen You, Renye Yan, Shiping Wen, Yingli Tian, Tingwen Huang
With the rise of Large Language Models (LLMs), significant progress has been made in research related to code automation within the field of software engineering. Among these developments, the code generation task has attracted considerable attention from both academia and industry due to its broad applicability. However, most existing studies focus primarily on improving the correctness of the code generated by LLMs, while overlooking the efficiency attributes of the code. As a result, LLMs often produce correct code without considering human preferences for efficiency. Efficient code not only conserves resources but also enhances the user experience. For online optimization scenarios, we find that the fixed clipping mechanism of Group Relative Policy Optimization (GRPO) can lead toentropy collapse, causing premature convergence to suboptimal solutions. To address this, we propose DC-GRPO. The method employs a multi-level reward that accounts for compilation success, functional correctness, and execution efficiency to alleviate sparse-reward issues. In addition, a dynamicclippingmechanism based on a four-quadrant analysis of advantage and generation probability adapts the clipping bounds via a nonlinear mapping, encouraging low-probability/high-reward responses while suppressing high-probability/low-reward ones. In our experiments on a 1.5B-parameter code model, DC-GRPO achieves an eff@5 of 13.4%, an improvement of 1.2 percentage points over standard GRPO, and shows consistent gains over other baselines on code-oriented LLMs. Our anonymous codes are available athttps://anonymous.4open.science/r/DC-GRPO-9E0D/.