arXiv:cs.LG· SiYuan Ma, Canran Xiao, Zikai Xiao, Albert Gao, Liang He, Xuan-Yu Wang, Shuying Cao, Xiaojun Jia·· 4 小时前AI 评分38
蛋白质语言模型中的线性适应度子空间实现样本高效定向进化
Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution
AI 导读
研究者提出 Linear Fitness Subspace(LFS)假设:突变引起的残基级表示变化中,一个紧凑且化验特异的方向集能让适应度变化从少量带标签变体中线性恢复,并据此构建 Subspace-Guided Evolutionary Search(SGES),在小样本上估计 LFS 后在子空间内完成代理建模、不确定性估计与采集。
正文
Abstract:Model-guided directed evolution seeks to identify high-fitness protein variants under limited oracle budgets. Protein language models (PLMs) provide rich representations for this task, but task-agnostic zero-shot scores can be misaligned with a target assay, while supervised search in high-dimensional embedding spaces can make surrogate modeling and uncertainty estimation sample-inefficient. We propose the Linear Fitness Subspace (LFS) hypothesis: within mutation-induced residue-level representation changes, a compact, assay-specific set of directions makes fitness variation linearly accessible from few labeled variants. This is a local, supervision-recoverable statement rather than a claim that protein fitness landscapes or global PLM geometry are universally linear. Building on this observation, we introduce Subspace-Guided Evolutionary Search (SGES), which estimates an LFS from a small initial sample and performs surrogate modeling, uncertainty estimation, and acquisition in the learned subspace. Across 10 core ProteinGym assays, 87 extended static-validation assays, and an 18-assay budgeted-search evaluation, SGES improves fitness prediction and search efficiency over zero-shot PLMs and recent ML-guided protein optimization baselines. Controlled comparisons with PCA, random projections, label-shuffled PLS, classical mutation features, and acquisition ablations further isolate the benefit of a fitness-aligned site-delta coordinate.
| Comments: | Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026 |
| Subjects: | Quantitative Methods (q-bio.QM); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07607 [q-bio.QM] |
| (or arXiv:2610.07607v1 [q-bio.QM] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07607 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: SiYuan Ma [view email]
[v1]
Tue, 6 Oct 2026 01:50:13 UTC (4,136 KB)
来源:arXiv:cs.LG · arxiv.org