跳到正文
arXiv:cs.LG· Xiukun Wei, Tian Xie, Ding Zhu, Xueru Zhang·· 4 小时前AI 评分44

自消耗生成模型与共同演化的人类偏好

Self-Consuming Generative Models with Co-Evolving Human Preferences

AI 导读

研究揭示,在完全依赖用户筛选合成数据的迭代训练中,初始偏见会被放大并导致系统收敛至多个单一均衡点之一;而注入参考数据可改变动态并产生唯一全局吸引均衡。作者据此提出一种算法,联合选择参考分布及其混合权重,以引导系统在最小化数据收集成本的同时保留期望属性。

正文

View PDF HTML (experimental)

Abstract:Generative models are increasingly trained in self-consuming iterative loops, where users curate preferred samples from model-generated candidates and the curated samples are used to train future generations of the model. Prior work has largely assumed fixed user preferences, but in practice exposure to model outputs gradually reshapes what users perceive as desirable, creating a feedback loop in which model distributions and user preferences co-evolve. We take a first step toward understanding the long-term behavior of such coupled dynamics. We show that when training relies entirely on user-curated synthetic data, iterative curation amplifies initial biases and drives the system toward one of multiple singleton equilibria in which the instance holding an initial advantage eventually dominates. In contrast, injecting reference data into training at a sufficiently large rate fundamentally changes the dynamics and yields a unique globally attracting equilibrium. Building on this insight, we study how reference-data injection can be used to control long-term outcomes, and propose an efficient algorithm that jointly selects a reference distribution and its mixing weight to steer the coupled system toward equilibria that preserve desired attributes while minimizing data collection costs.
Comments: Published as a conference paper at NeurIPS 2026
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.09415 [cs.LG]
  (or arXiv:2610.09415v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09415

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xiukun Wei [view email]
[v1] Wed, 7 Oct 2026 04:18:20 UTC (5,555 KB)

来源:arXiv:cs.LG · arxiv.org