科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ ACM Transactions on Software Engineering and Methodology2025-10-27· Deed

Exploring Data-Efficient Adaptation of Large Language Models for Code Generation

Xue Jiang, Yihong Dong, Zhiyuan Fan, Zhi Jin, Wenpin Jiao, Ge Li

原始摘要(英文原文)· Original abstract
Although Large Language Models (LLMs) have made significant progress in code generation, they still struggle with code generation tasks in specific scenarios. These scenarios usually necessitate the adaptation of LLMs to fulfill specific needs, but the limited training data available in practice leads to poor code generation performance. Therefore, how to effectively adapt LLMs to new scenarios with few training data is a major challenge for current code generation. In this article, we propose a novel adaptation approach named Data-Efficient adaptation with Error-Driven Learning (DEED) for code generation. D EED leverages the errors made by LLMs as learning opportunities, using error revision to overcome their own shortcomings, thus achieving efficient learning. Specifically, D EED involves identifying error code generated by LLMs, employing S elf-Revise for code revision, optimizing the model with revised code, and iteratively adapting the process for continuous improvement. Experimental results show that, compared to other mainstream fine-tuning approaches, D EED achieves superior performance with few training data, showing an average relative improvement of 46.2% in Pass@1 on multiple code generation benchmarks. We also validate the effectiveness of S elf-Revise, which generates revised code that optimizes the model more efficiently compared to the code samples from datasets. Moreover, D EED consistently demonstrates strong performance across various LLMs, underscoring its applicability.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

Exploring Data-Efficient Adaptation of Large Language Models for Code Generation — 科研速览 Science Skim