arXiv:cs.LG· Jeremy Levy, Ariel Larey, Yury Nahshan, Raizy Kellerman, Elay Dahan, Amit Bleiweiss, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Marissa Wirth, Simon Lee, Dung Hoang, Noam D. Beckmann, Shane O'Connell, Nicole Bussola, Alexander W. Charney, Yoli Shavit, Nati Daniel·· 2 天前AI 评分35
VANDAM:用 DNA 分子先验增强基因组基础模型训练
VANDAM: Viewing a nucleotide sequence with DNA molecular priors
AI 导读
VANDAM 是一个为基因组基础模型(GFM)引入 DNA 分子先验的训练框架,在自监督训练中从池化表示预测区域分子属性,并在有功能标签时于输入端注入局部特征。该框架在四类架构家族、九项留出基因组任务上持续提升下游性能,探针实验显示分子先验可泛化到未见过的分子属性。
正文
Authors:Jeremy Levy, Ariel Larey, Yury Nahshan, Raizy Kellerman, Elay Dahan, Amit Bleiweiss, Guy Leib, Omri Nayshool, Dan Ofer, Tal Zinger, Dan Dominissini, Gideon Rechavi, Marissa Wirth, Simon Lee, Dung Hoang, Noam D. Beckmann, Shane O'Connell, Nicole Bussola, Alexander W. Charney, Yoli Shavit, Nati Daniel
Abstract:Contemporary Genomic Foundation Models (GFMs) rely on a DNA-as-a-string paradigm that employs masked token prediction objectives for pretraining. However, this abstraction does not explicitly model the biochemical, structural, and physical properties essential to biological function. Many molecular properties can be estimated from sequence using established biophysical models, so their utility lies not in providing an independent modality, but in introducing priors that training objectives can explicitly exploit. We introduce VANDAM, a framework that extends the training of GFMs with DNA molecular priors. In self-supervised training, VANDAM predicts regional molecular properties from pooled representations. When functional labels are available and can reward retaining molecular priors, local features are additionally injected at the input. VANDAM consistently improves downstream performance across four architecture families and nine held-out genomic tasks by complementing token-based objectives. Probing experiments further demonstrate that the use of molecular priors generalizes to other unseen molecular properties.
| Comments: | 25 pages, 3 figures, including appendices |
| Subjects: | Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.00411 [cs.LG] |
| (or arXiv:2610.00411v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00411 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jeremy Levy Mr [view email]
[v1]
Wed, 30 Sep 2026 14:03:57 UTC (828 KB)
来源:arXiv:cs.LG · arxiv.org