跳到正文
arXiv:cs.LG· Julian Kleutgens, Mauricio Tec, Claudio Battiloro, Francesca Dominici, Giannis Daras·· 3 小时前

RefineMix:在数据稀缺时用分布外数据训练离散扩散模型

Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient Learning

AI 导读

RefineMix 框架通过在选择性的扩散时间步引入分布外数据来训练离散扩散模型,在严重数据稀缺条件下提升泛化能力且不使采样分布产生偏差。在蛋白质序列生成任务中,仅用 197 个域内样本微调,生成的蛋白质同时满足新颖、可折叠且属于同一家族的比例相比标准微调接近翻倍。

正文

View PDF HTML (experimental)

Abstract:We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications. RefineMix uses out-of-distribution data at selected diffusion times to improve generalization without biasing the sampling distribution. Although this strategy has been explored in continuous diffusion, discrete diffusion presents a distinct challenge: unlike Gaussian noise, masking preserves domain information in surviving tokens, limiting the use of related data at high noise levels. At low noise levels, however, the domains effectively disjoint supports become an advantage, allowing the model to learn from both in-domain and out-of-distribution data without biasing the sampler. We formalize these intuitions and provide a theoretical analysis for the proposed method. Experimentally, across five domain-shift settings, RefineMix matches or outperforms in-domain finetuning and data mixing. For protein sequence generation, finetuning with just 197 in-domain examples nearly doubles the fraction of generated proteins that are simultaneously novel, foldable, and in-family compared to standard finetuning.
Comments: 10 pages. Accepted at NeurIPS 2026 Workshops (BeNTo, DiffuLM)
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.12340 [cs.LG]
  (or arXiv:2610.12340v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.12340

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Julian Kleutgens [view email]
[v1] Thu, 8 Oct 2026 17:15:19 UTC (517 KB)

来源:arXiv:cs.LG · arxiv.org