arXiv:cs.AI· Yuanhao Ban, I-Hung Hsu, Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh·· 4 小时前AI 评分44
组合偏好奖励与评分规则奖励,后训练前沿文生图模型
Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards
AI 导读
研究者提出一种组合互补奖励信号的后训练方案,用于开放域文生图:偏好奖励基于大规模人类偏好数据以 Bradley-Terry 目标训练,评分规则奖励评估提示词忠实度并防止奖励攻击。
正文
Abstract:Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference. In this work, we develop a simple and effective post-training recipe for open-domain text-to-image generation based on the composition of complementary reward signals. Our reward system consists of two main components: a preference reward, trained on large-scale human preference data using a Bradley-Terry objective to capture overall human aesthetic and perceptual preferences, and rubric-based rewards, which explicitly evaluate prompt faithfulness and other desirable properties while providing safeguards against reward hacking. A key challenge is how to combine these heterogeneous reward signals. We show that a naive weighted average leads to suboptimal optimization behavior, and propose a simple reward composition strategy that more effectively balances preference optimization with rubric satisfaction. In the Arena text-to-image leaderboard (this https URL), our RL-trained Flux2dev achieves an Elo rating 69 points above the base model, and our post-trained Ideogram-4 surpasses every open-source model on the leaderboard, reaching an Elo of 1223.5. (Claims of state-of-the-art performance are based on the Arena leaderboard snapshot as of September 4, 2026.) Our results suggest that effective rewards for frontier generative-model training require broad coverage of user intent and robustness to exploitation under optimization. To support reproducible research, we release Arena-T2I-Training, a 1K subset of training data that recovers some gains of full-scale training, providing a resource that we hope will facilitate future work on post-training for text-to-image models.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02967 [cs.CV] |
| (or arXiv:2610.02967v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02967 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yuanhao Ban [view email]
[v1]
Fri, 2 Oct 2026 08:05:34 UTC (5,159 KB)
来源:arXiv:cs.AI · arxiv.org