Low-Rank Contextual Reinforcement Learning from Heterogeneous Human Feedback
成果类型:
Article
署名作者:
Lee, Seong Jin; Sun, Will Wei; Liu, Yufeng
署名单位:
University of North Carolina; University of North Carolina Chapel Hill; Purdue University System; Purdue University; University of Michigan System; University of Michigan
刊物名称:
JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION
ISSN/ISSBN:
0162-1459; 1537-274X
DOI:
10.1080/01621459.2026.2674413
发表日期:
2026-04-03
页码:
1037-1050
关键词:
Distribution Shift
Large language models
Low-rankness
Offline Reinforcement Learning
摘要:
Reinforcement learning from human feedback (RLHF) has become a cornerstone for aligning large language models with human preferences. However, the heterogeneity of human feedback, driven by diverse individual contexts and preferences, poses significant challenges for reward learning. To address this, we propose a Low-rank Contextual RLHF (LoCo-RLHF) framework that integrates contextual information to better model heterogeneous feedback while maintaining computational efficiency. Our approach builds on a contextual preference model, leveraging the intrinsic low-rank structure of the interaction between user contexts and query-answer pairs to mitigate the high dimensionality of feature representations. Furthermore, we address the challenge of distributional shifts in feedback through our Pessimism in Reduced Subspace (PRS) policy, inspired by pessimistic offline reinforcement learning techniques. We theoretically demonstrate that our policy achieves a tighter sub-optimality gap compared to existing methods. Extensive experiments, ranging from synthetic simulations to an analysis of the real-world PersonalLLM benchmark, validate the effectiveness of LoCo-RLHF and demonstrate its superior performance in personalized settings. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
来源URL: