跳到正文
arXiv:cs.LG· Jae Won Choi, Ryoonki Hong, Alan Liang, Manjula Adiveppa Wader, Bingsong Zeng, Peiyang Tang, Longwei Liu, Ruishan Liu·· 4 小时前AI 评分30

PIT-GCL:用拓扑图对比学习预测蛋白质相互作用

PIT-GCL: Protein Interaction using Topological Graph Contrastive Learning

AI 导读

PIT-GCL 是一个双塔结构感知框架,从氨基酸序列、Cα 点云和全局持续同调描述符独立编码每个蛋白质,通过结构感知 Transformer 与双向交叉注意力实现潜在空间软对接,并以二元交叉熵加 NT-Xent 对比目标训练。

正文

View PDF HTML (experimental)

Abstract:Protein binding prediction is central to target identification, therapeutic binder design, and large scale screening, yet remains challenging because binding depends on sequence, three dimensional geometry, and global structural organization. Recent folding models such as AlphaFold3 and Boltz-2 have substantially improved structure prediction, but their confidence outputs (pLDDT, pTM, ipTM) are not specifically designed for binary binding prediction, and dedicated structure aware predictors often require bound complex structures that are unavailable at screening scale. We introduce PIT-GCL, a dual tower structure aware framework that encodes each protein independently from its amino acid sequence, C{\alpha} point cloud, and a global persistent homology descriptor. Each tower combines residue ESM-2 embeddings with a topological summary computed from the H0 and H1 persistence landscapes of a Vietoris-Rips filtration, and processes the resulting tokens with a structure aware Transformer in which pairwise C{\alpha} distances enter as a learned attention bias. A bidirectional cross attention module then performs latent space soft docking between the two per-protein representations, and the model is trained with a combined binary cross entropy and NT-Xent contrastive objective. On three binary interaction prediction benchmarks, general PPI on PPIRef, TCRpMHC binding on STAG, and whole chain pairs on PPB-Affinity, PIT-GCL outperforms representative sequence based, structure aware, and task specific baselines on general PPI under our evaluation, and is the only method above chance on PPB-Affinity; on TCR-pMHC it leads at a fixed decision threshold but is outranked by a task specific sequence model. Because each protein is encoded independently in the first phase, its representation can be precomputed and reused across candidate pairs, which is convenient for large scale screening.
Comments: 12pages, 6figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.04850 [cs.LG]
  (or arXiv:2610.04850v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.04850

arXiv-issued DOI via DataCite

Submission history

From: Jae Choi [view email]
[v1] Sun, 4 Oct 2026 01:30:42 UTC (113 KB)
[v2] Tue, 6 Oct 2026 21:57:50 UTC (113 KB)

来源:arXiv:cs.LG · arxiv.org