arXiv:cs.LG· Christina Hahn, Shangbin Feng, Dean Light, Swastik Roy, Hila Gonen, Yulia Tsvetkov·· 6 小时前AI 评分38
基于 Stackelberg 博弈的多 LLM 协同对齐
Multi-LLM Collaborative Alignment via Stackelberg Games
AI 导读
研究者提出 Stackelberg Alignment,一种受博弈论启发的 leader-follower 框架,将指令选择转化为自适应课程,由 EXP3 bandit 作为 leader 分配采样预算,语言模型作为 follower 通过 DPO 或 GRPO 学习偏好信号。
正文
Abstract:A pool of language models can collaborate and improve collectively by learning from one another's responses. These interactions depend on the instructions used during training. Existing methods typically sample instructions uniformly, even though their usefulness may change as the models improve: an instruction on which models' responses once differed in quality may later be answered equally well, while a previously difficult instruction may begin to provide a useful learning signal. We propose Stackelberg Alignment, a game-theory-inspired leader-follower framework that turns instruction selection into an adaptive curriculum. An EXP3 bandit acts as the leader, allocating a fixed sampling budget across instructions and updating its sampling distribution using a reward that combines instruction difficulty and response discriminability. The language models act as followers: they respond to the selected instructions, evaluate one another's responses, and learn from the resulting preference signals through DPO or GRPO. The framework uses Elo-style reputation-weighted peer judgment and reputation-based opponent matching to support reliable and competitive model interactions. Experiments across three heterogeneous model pools and 12 benchmarks spanning scientific discovery, reasoning, code, instruction following, and knowledge show that Stackelberg Alignment achieves the highest macro-average across three diverse model pools, outperforming the strongest training-time baseline by up to 7.4% and the best static inference baseline by 12-25%. Analysis confirms that the adaptive leader concentrates duels on the most informative instructions, and ablations show that both reputation-weighted judgment and reputation-based matching improve the effectiveness of multi-LLM evolution.
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.39076 [cs.AI] |
| (or arXiv:2609.39076v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.39076 arXiv-issued DOI via DataCite |
Submission history
From: Shangbin Feng [view email]
[v1]
Wed, 30 Sep 2026 06:13:56 UTC (981 KB)
[v2]
Tue, 6 Oct 2026 23:41:48 UTC (981 KB)
来源:arXiv:cs.LG · arxiv.org