arXiv:cs.LG· Tianlong Nan, Xiaopeng Li, Christian Kroer, Tianyi Lin·· 5 小时前AI 评分35
面向迭代式 Nash 偏好优化的高效探索方法 ENPO
Efficient Exploration for Iterative Nash Preference Optimization
AI 导读
针对迭代式 Nash 学习(NLHF)中探索不足的问题,研究者提出 Exploratory Nash Preference Optimization(ENPO),通过结合 SFT 型正则化与对抗式策略探索,消除了标准迭代 NLHF 对逆 KL 正则化参数的指数依赖,且无需 minimax oracle 或显式偏好模型估计。
正文
Abstract:Preference alignment is central to improving large language models (LLMs), but reward-based formulations can be restrictive when human preferences are non-transitive. Nash learning from human feedback (NLHF) addresses this limitation by modeling alignment as a preference game and seeking a Nash equilibrium. However, the learning-theoretic foundations of scalable NLHF remain limited: existing regret guarantees rely on explicit preference-model estimation and minimax oracles, whereas simpler iterative methods lack such guarantees. We study online iterative NLHF and identify exploration as a key obstacle. First, we show that standard iterative NLHF can incur an exponential dependence on the inverse KL-regularization parameter, demonstrating that implicit exploration through policy updates can be insufficient. We then propose Exploratory Nash Preference Optimization (ENPO), which combines a SFT-type regularization with adversarial policy exploration. ENPO eliminates this exponential dependence without requiring minimax oracles or explicit preference-model estimation. We further introduce Bonus-Explorer ENPO (BENPO), which uses additional oracles to achieve an $O(\log T)$ regret bound. Finally, we develop Direct ENPO (DENPO), a practical variant of ENPO for fine-tuning LLMs. Experiments with Llama-3-8B-Instruct demonstrate consistent improvements over the evaluated RLHF and NLHF baselines across multiple benchmarks.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.01382 [cs.LG] |
| (or arXiv:2606.01382v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2606.01382 arXiv-issued DOI via DataCite |
Submission history
From: Tianlong Nan [view email]
[v1]
Sun, 31 May 2026 18:11:26 UTC (93 KB)
[v2]
Fri, 2 Oct 2026 10:57:59 UTC (94 KB)
来源:arXiv:cs.LG · arxiv.org