arXiv:cs.AI· Miruna Cretu, Alex Abrudan, Antonia Panescu, Tynan Perez, Rishabh Anand, N. Benjamin Erichson, Michael W. Mahoney, Samuel Blau, Joseph Jacobson, Rafael G\'omez-Bombarelli, Rex Ying, Tuomas Knowles, Pietro Li\`o, Alex Morehead·· 4 小时前
Zatom-2:跨领域生成建模的原子级数据多任务预训练
Zatom-2: Multitask Pretraining on Atomistic Data for Generative Modeling across Domains
AI 导读
Zatom-2 是一个在 OMol25 和 OMat24 电子结构数据集约 500 万个结构上预训练的原子级生成模型,采用多尺度 Transformer 架构结合条件流匹配,支持力条件控制和生成、结构预测等预训练任务。
正文
Authors:Miruna Cretu, Alex Abrudan, Antonia Panescu, Tynan Perez, Rishabh Anand, N. Benjamin Erichson, Michael W. Mahoney, Samuel Blau, Joseph Jacobson, Rafael Gómez-Bombarelli, Rex Ying, Tuomas Knowles, Pietro Liò, Alex Morehead
Abstract:Unified atomistic modeling has the potential to accelerate discovery in chemistry, materials science, and biology by bridging data-rich chemical domains and data-scarce biological contexts. However, existing generative approaches to atomistic modeling remain highly specialized to scientific disciplines (chemistry vs. biology) or do not leverage both high-volume organic (molecule) and inorganic (material) data for general-purpose pretraining. To this end, we introduce Zatom-2, an atomistic generative model pretrained on approximately five million structures from the OMol25 and OMat24 electronic structure datasets. Zatom-2 features a multiscale Transformer architecture coupled with conditional flow matching that supports force conditioning and foundational pretraining tasks such as generation, structure prediction, and prediction of molecular and material energies and forces. Empirically, Zatom-2 achieves better molecular distribution fidelity than Zatom-1 and achieves strong performance on existing molecule and material generation benchmarks. Zatom-2 demonstrates the ability to control sample generation across low- and high-force regimes, and enhances protein generation in a low-data setting through joint generative-predictive pretraining and transfer learning, increasing protein backbone designability in a length extrapolation setting from 67.8% without pretraining to 74.8% after finetuning on 2,000 protein domains.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11454 [cs.LG] |
| (or arXiv:2610.11454v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11454 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: N. Benjamin Erichson [view email]
[v1]
Thu, 8 Oct 2026 08:09:17 UTC (11,304 KB)
来源:arXiv:cs.AI · arxiv.org