arXiv:cs.AI(全量分类)· Shihab Ahmed, Debamita Ghosh, David Tang, Yudan Wang, Alvaro Velasquez, Yue Wang·· 5 小时前AI 评分44
Robust Nash Alignment:面向偏好不确定性的鲁棒纳什对齐框架
Robust Nash Alignment under Preference Uncertainty
AI 导读
研究者提出 Robust Nash Alignment,一个针对不确定成对偏好的博弈论对齐框架,让主学习者在对抗竞争者与偏好模糊集下最大化最差胜率,并给出最差性能的认证下界。
正文
Abstract:Preference-based alignment methods typically optimize against a single preference model, and can therefore be brittle when pairwise preferences are uncertain: noisy, heterogeneous, or shift after deployment. To address these issues, we propose Robust Nash Alignment, a game-theoretic framework for alignment to uncertain pairwise preferences. Our formulation has a major learner seeking a policy with a large worst-case win rate against both an adversarial competitor and any preference kernel lying in an ambiguity set around a nominal preference. When the ambiguity set captures the uncertainty in preferences, the resulting robust objective of the game directly yields a certified lower bound on worst-case performance. However, we note this problem is computationally challenging to optimize, and to address this, we introduce a four-player primal-dual proxy game involving the leader policy, follower policy, adversarial kernel, and dual variable, and develop a single-loop optimistic mirror descent-ascent algorithm for it. We show that the proxy always lower-bounds the truncated hard-constrained objective, quantify the proxy-to-hard gap, and characterize an exactness condition under which the proxy recovers the robust objective. We then prove an \(\mathcal{O}(1/\sqrt{T})\) average-iteration convergence for the proxy-game duality gap, which implies a near-optimal robust policy for the original robust objective. Experiments on controlled tabular games and LLM alignment with uncertain preference further validate the convergence theory and show improved performance over nominal baselines.
| Comments: | 38 pages, accepted at 2026 40th Advances in Neural Information Processing System (NeurIPS) |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.00715 [cs.AI] |
| (or arXiv:2610.00715v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00715 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shihab Ahmed [view email]
[v1]
Wed, 30 Sep 2026 21:05:43 UTC (533 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org