arXiv:cs.LG(机器学习,全量分类)· Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff·· 14 小时前AI 评分42
IDiom:用稀疏自编码器特征强化生成内在无序蛋白区域
Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
AI 导读
研究团队提出 IDiom,一个在 5400 万条 AlphaFold Database 预测 IDR 数据集 IDiom-DB 上训练的自回归蛋白质语言模型,并配套 RL-SAE 后训练方法,通过奖励激活指定特征集的序列来控制功能相关序列模式。
正文
Abstract:Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcement learning with sparse autoencoder features (RL-SAE), a post-training method that rewards the generation of sequences that activate specified feature sets. Across eight IDR design tasks, RL-SAE sequences activate, on average, 90% of 30 targeted features, compared to 24% for activation steering. We demonstrate that RL-SAE improves the predicted subcellular localization and transcriptional activity of generated IDRs compared to steering and supervised fine-tuning, and enables features associated with distinct biological functions to be combined within individual sequences. Thus, IDiom and RL-SAE enable interpretable and composable IDR design through explicit control of function-associated sequence features. More broadly, RL-SAE could extend to other protein design settings where interpretable features provide useful design targets. Code is available at this https URL.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02189 [cs.LG] |
| (or arXiv:2610.02189v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02189 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jason Liu [view email]
[v1]
Thu, 1 Oct 2026 17:59:20 UTC (5,830 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org