Ruili Jiang, Kehai Chen, Xuefeng Bai, Z. H. He, Juntao Li, Muyun Yang, Tiejun Zhao, Liqiang Nie, Min Zhang
The recent surge in versatile large language models (LLMs) demonstrates remarkable success across a wide range of contexts. A key factor contributing to this success is LLM alignment, in which human preference learning plays a decisive role in steering the models’ capabilities toward fulfilling human objectives. In this survey, we review the progress in human preference learning within a unified framework, aiming to provide a comprehensive perspective on established methodologies while exploring avenues to further advance LLM alignment. Specifically, we categorize human preference feedback based on data sources and formats, summarize techniques for human preference modeling and usage, and present an overview of prevailing evaluation protocols for LLM alignment. Finally, we discuss the existing challenges and identify potential directions for future research, with a particular emphasis on generalizability, transferability, and controllability.