跳到正文
arXiv:cs.LG· Kasper K. Jakobsen, Eli N. Weinstein·· 7 小时前AI 评分40

信息密集型合成用于分子发现

Information-Dense Synthesis for Molecular Discovery

AI 导读

研究者提出一种算法控制的随机合成方法,通过设计并合成复杂混合物、以池为单位测试再解卷积分子-活性图谱,来高效搜索大范围分子空间。理论上该方法可将从 d 个候选分子中寻找最优分子所需的实验次数从 O(d) 降至 O(log d) 或 O(1);在估计的蛋白质适应度景观模拟中,其找到活性分子所需实验次数比现有贝叶斯优化方法少一个数量级。

正文

View PDF HTML (experimental)

Abstract:Machine learning can accelerate molecular discovery by designing molecules and planning experiments. However, many scientific challenges demand molecules with very rare properties, and in this sparse setting, existing algorithms offer little gain over random guessing. We propose a method to efficiently search large regions of molecular space using algorithmically controlled stochastic synthesis. Rather than design, make and test individual molecules, we design and make complex mixtures, test them as a pool, then deconvolute the molecule-activity map. We optimize synthesis to encode maximal information. Theoretically, this approach can reduce the number of experiments required to find the optimal molecule among $d$ candidates from $\mathcal{O}(d)$ to $\mathcal{O}(\log d)$ or $\mathcal{O}(1)$. In simulation, on estimated protein fitness landscapes, it finds active molecules with an order of magnitude fewer experiments than existing Bayesian optimization methods.
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Chemical Physics (physics.chem-ph); Biomolecules (q-bio.BM)
Cite as: arXiv:2610.08495 [stat.ML]
  (or arXiv:2610.08495v1 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2610.08495

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Kasper Krunderup Jakobsen [view email]
[v1] Tue, 6 Oct 2026 15:04:49 UTC (2,138 KB)

来源:arXiv:cs.LG · arxiv.org