On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization

成果类型:
Article
署名作者:
Xiao, Jiancong; Li, Ziniu; Xie, Xingyu; Getzen, Emily; Fang, Cong; Long, Qi; Su, Weijie J.
署名单位:
University of Pennsylvania; Pennsylvania Medicine; The Chinese University of Hong Kong, Shenzhen; National University of Singapore; Peking University; University of Pennsylvania
刊物名称:
JOURNAL OF THE AMERICAN STATISTICAL ASSOCIATION
ISSN/ISSBN:
0162-1459; 1537-274X
DOI:
10.1080/01621459.2025.2555067
发表日期:
2025-10-02
页码:
2154-2164
关键词:
large language model Preference matching Reinforcement learning from human feedback Reward model
摘要:
Accurately aligning large language models (LLMs) with human preferences is crucial for informing fair, economically sound, and statistically efficient decision-making processes. However, we argue that the predominant approach for aligning LLMs with human preferences through a reward model-reinforcement learning from human feedback (RLHF)-suffers from an inherent algorithmic bias due to its Kullback-Leibler-based regularization in optimization. In extreme cases, this bias could lead to a phenomenon we term preference collapse, where minority preferences are virtually disregarded. To mitigate this algorithmic bias, we introduce preference matching (PM) RLHF, a novel approach that provably aligns LLMs with the preference distribution of the reward model under the Bradley-Terry-Luce/Plackett-Luce model. Central to our approach is a PM regularizer that takes the form of the negative logarithm of the LLM's policy probability distribution over responses, which helps the LLM balance response diversification and reward maximization. Notably, we obtain this regularizer by solving an ordinary differential equation that is necessary for the PM property. For practical implementation, we introduce a conditional variant of PM RLHF that is tailored to natural language generation. Finally, we empirically validate the effectiveness of conditional PM RLHF through experiments on the OPT and Llama-family models, demonstrating a 29%-41% improvement in alignment with human preferences, as measured by a certain metric, compared to standard RLHF. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
来源URL: