跳到正文
arXiv:cs.LG· Minjae Chung, Clara Li, Malar Paavai Muthukumaran, Shaunna Wang, Aniket Ramkrishnan Iyer, Shaun Qien Yeau Tan, Harinishree Sathu, Micky C. Nnamdi, J. Ben Tamo, Benoit Louis Marteau, May Dongmei Wang·· 3 小时前AI 评分40

超越随机划分:在化学与生物学驱动的分布偏移下评估药物-靶点亲和力模型

Beyond Random Splits: Evaluating Drug-Target Affinity Models Under Chemically and Biologically Motivated Distribution Shifts Copy

AI 导读

研究从 ChEMBL 和 BindingDB 整理 718,800 个药物-蛋白对,比较 Morgan 指纹+protein-CNN 基线及 12 种架构(4 种药物表示 × 3 种 ESM-2 交互模式)。

正文

Authors:Minjae Chung, Clara Li, Malar Paavai Muthukumaran, Shaunna Wang, Aniket Ramkrishnan Iyer, Shaun Qien Yeau Tan, Harinishree Sathu, Micky C. Nnamdi, J. Ben Tamo, Benoit Louis Marteau, May Dongmei Wang

View PDF HTML (experimental)

Abstract:Drug-target affinity (DTA) prediction is widely used to prioritize candidate compounds before costly experimental screening. DTA models are often compared under a single data split, even though deployment may require extrapolation to new chemical series, new protein targets, or both. We ask whether the distribution shift used for evaluation changes which architecture appears best. We curate 718,800 unique drug-protein pairs from the ChEMBL and BindingDB datasets. We compare a Morgan-fingerprint + protein-CNN baseline with 12 controlled architectures that combine four drug representations with three ESM-2 interaction modes. Mean validation RMSE increases from 0.950 and 0.945 under scaffold and fingerprint-cluster OOD to 1.299 and 1.321 under protein-cluster and dual OOD. Model rankings are similar across the two chemical shifts (tau = 0.79), but agreement with scaffold OOD falls under protein OOD (tau = 0.39) and reverses under dual OOD (tau = -0.55). Held-out evaluation, repeated seeds, group-aware bootstrap analysis, and a size-matched control support the same conclusion: architecture selection depends on the form of extrapolation, not only on average error or training-set size. DTA benchmarks should therefore match the chemical and target shifts expected at deployment.
Comments: Accepted to the NeurIPS 2026 Workshop on AI for Drug Discovery (AI4DD)
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.03456 [cs.LG]
  (or arXiv:2610.03456v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03456

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Minjae Chung [view email]
[v1] Fri, 2 Oct 2026 15:33:20 UTC (1,728 KB)

来源:arXiv:cs.LG · arxiv.org