arXiv:cs.LG· Loka Li, Jin Tian, Kun Zhang·· 4 小时前AI 评分35
CausalBind:面向蛋白质-分子虚拟筛选的因果建模与学习
CausalBind: Causal Modeling and Learning for Protein-Molecule Virtual Screening
AI 导读
NeurIPS 2026 Oral 论文提出 CausalBind,用 V-structure 因果模型和 Heckman 式选择形式化蛋白质-分子结合中的稀疏跨模态交互,并证明在结构稀疏条件下交互概念可成分级识别。该方法在 DUD-E 和 LIT-PCBA 基准上一致优于强检索基线,LIT-PCBA 早期富集提升最大,并在靶点和骨架级分布外划分上具备泛化能力,代码已开源。
正文
Abstract:Protein-molecule virtual screening is increasingly cast as a problem of representation learning in a shared embedding space. Existing methods rely on dense holistic alignment, entangling invariant binding determinants with nuisance correlations and limiting transfer to new targets. It has been noted that binding in protein-molecule systems involves sparse cross-modality interactions: binding is governed by a small contact interface and a few decisive local interactions (e.g., hydrogen bonds, hydrophobic contacts, and salt bridges) rather than the global structures of the protein and molecule. We hypothesize that uncovering and leveraging sparse interaction patterns is critical for generalization beyond the training data, as these patterns are reusable and expected to improve performance across different scenarios. In this paper, we aim to identify and leverage sparse interaction patterns, and verify our hypothesis. Since the training data contain only observed binding pairs, we formalize this prior via a V-structure causal model under Heckman-style selection, and establish three theoretical results: (i) the latent concepts of interacting proteins and molecules are not identifiable without appropriate sparsity constraints; (ii) these concepts and their sparse interactions are component-wise identifiable under structural sparsity conditions; and (iii) a low-rank relaxation of these conditions yields subspace identifiability of the concepts and interactions. Inspired by these principles, we propose CausalBind with three implementation variants. Extensive experiments on DUD-E and LIT-PCBA benchmarks show that all variants consistently outperform strong retrieval baselines, with the largest gains on LIT-PCBA early enrichment, and further generalize to target- and scaffold-level out-of-distribution splits. Code is available at this https URL.
| Comments: | NeurIPS 2026 (Oral) |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07340 [cs.LG] |
| (or arXiv:2610.07340v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07340 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Loka Li [view email]
[v1]
Mon, 5 Oct 2026 20:14:51 UTC (667 KB)
来源:arXiv:cs.LG · arxiv.org