跳到正文
arXiv:cs.LG· T\'elio Cropsal, Roc\'io Mercado·· 4 小时前AI 评分37

分子编码器能否在没有实验读出数据的情况下预测表型活性?

Can phenotypic activity be predicted without experimental readouts?

AI 导读

研究评估了在配对分子-形态数据上对比预训练的分子编码器(如 CLOOME 和 CellCLIP)能否作为表型预测的廉价替代方案,从而免去运行 Cell Painting 实验。在两项 Cell Painting 筛选中测试六种表征后,一旦控制了编码器自身预训练边界的泄露以及表型活性与细胞毒性之间的相关性这两个混杂因素,预训练分子编码器相较普通理化描述符并无明显优势,且毒性通常比表型活性更易预测。

正文

View PDF HTML (experimental)

Abstract:Molecular encoders contrastively pretrained on paired molecule-morphology data, such as CLOOME and CellCLIP, have been proposed as cheap surrogates for phenotypic prediction, avoiding the need to run a Cell Painting assay. We evaluate this idea for these molecular encoders under a protocol designed to control for two confounds that can inflate apparent performance: leakage across an encoder's own pretraining boundary, and the correlation between phenotypic activity and cytotoxicity. Testing six representations, including a non-pretrained MLP control matching CLOOME's input and layer count, on two distinct Cell Painting screens, we find that once these confounds are controlled for, the pretrained molecular encoders show no clear advantage over plain physicochemical descriptors, and that toxicity is generally easier to predict than phenotypic activity across representations. Our results suggest leakage-aware, confound-controlled evaluation should be standard practice before phenotype-pretrained encoders are trusted as surrogates for phenotypic drug discovery.
Comments: Accepted to the ML4Molecules: Agentic Systems for Molecular Sciences Workshop at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026)
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.07997 [cs.LG]
  (or arXiv:2610.07997v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07997

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Télio Cropsal [view email]
[v1] Tue, 6 Oct 2026 08:56:23 UTC (94 KB)

来源:arXiv:cs.LG · arxiv.org