arXiv:cs.LG· Parthasarathy Suryanarayanan, Susanta Das, Shreyans Sethi, Kenneth M. Merz, Jr., Joseph A. Morrone·· 5 小时前AI 评分34
GRACE:通过早期融合的几何残差加合物条件化预测碰撞截面
Predicting Collision Cross Sections with GRACE: Geometric Residual Adduct Conditioning via Early-fusion
AI 导读
GRACE 是一种 3D 碰撞截面(CCS)预测器,通过早期融合的几何残差加合物条件化来适配预训练分子几何编码器。在超过 9,000 条实验分子-加合物 CCS 记录上,GRACE 在随机、scaffold 和加合物敏感三种划分上分别取得 1.67%、2.11% 和 2.36% 的平均百分比差异,均为所评估学习模型中最佳。
正文
Abstract:Collision cross section (CCS), derived from ion mobility mass spectrometry, is a common descriptor for molecular annotation. Prediction is challenging for machine learning models because it reflects the size, shape, and ionization state of a gas-phase molecular ion. Most predictors either ignore explicit 3D structure or treat adduct identity as a late categorical feature, which limits their ability to capture adduct-dependent geometric effects. We present GRACE (Geometric Residual Adduct Conditioning via Early-fusion), a 3D CCS predictor that adapts a pretrained molecular geometry encoder using geometric residual adduct conditioning via early fusion. GRACE combines two inductive biases: a residual objective relative to an adduct-aware physical descriptor baseline and adduct conditioning within the encoder via a learned adduct token and low-rank attention adapters. We evaluate the model on a curated set of over 9,000 experimental molecule-adduct CCS records with random, scaffold, and adduct-sensitive splits designed to separate interpolation, scaffold generalization, and adduct-driven generalization. GRACE achieves the best mean percentage difference among the evaluated learned models on all three splits: 1.67% on the random split, 2.11% on the scaffold split, and 2.36% on the adduct-sensitive split. Diagnostic analyses suggest that residual learning stabilizes training by removing the dominant mass-CCS trend, while early fusion improves adduct-sensitive prediction relative to late fusion. Across four independent external test sets, GRACE shows consistently lower error than the other evaluated models. On a held-out set, GRACE also attains the lowest mean percent difference when compared with four previously reported physics-based workflows. These results support residual learning and encoder-level adduct conditioning as practical inductive biases for fast, accurate CCS prediction.
| Comments: | 35 pages, including 10 figures and 18 tables |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Biomolecules (q-bio.BM) |
| Cite as: | arXiv:2609.12223 [cs.LG] |
| (or arXiv:2609.12223v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.12223 arXiv-issued DOI via DataCite |
Submission history
From: Joseph Morrone [view email]
[v1]
Thu, 10 Sep 2026 21:31:11 UTC (708 KB)
[v2]
Thu, 1 Oct 2026 23:09:36 UTC (709 KB)
来源:arXiv:cs.LG · arxiv.org