Matteo Cercola, Valeria Capretti, Simone Formentin
Learning from human preferences is essential for aligning machine learning models with subjective judgments, but collecting preference data is costly. We study a hybrid framework that combines the scalability of Reinforcement Learning from Human Feedback (RLHF), which trains neural reward models from pairwise comparisons, with the sample efficiency of Preferential Bayesian Optimization. The method integrates Laplace-based Bayesian uncertainty estimation to guide informative preference queries. On high-dimensional Rosenbrock optimization, the approach successfully converges in problems with up to 50 dimensions. In large language model (LLM) fine-tuning, it improves reward-model accuracy by a value within 6 − − 14 % under limited annotation budgets.