跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Jason X. Liu, Sebastian Ibarraran, Frank Hu, Soojung Yang, Xinyu A. Feng, Abigail Park, Anagha Aneesh, Lacramioara Bintu, Alexander R. Dunn, Grant M. Rotskoff·· 14 小时前AI 评分42

IDiom:用稀疏自编码器特征强化生成内在无序蛋白区域

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

AI 导读

研究团队提出 IDiom,一个在 5400 万条 AlphaFold Database 预测 IDR 数据集 IDiom-DB 上训练的自回归蛋白质语言模型,并配套 RL-SAE 后训练方法,通过奖励激活指定特征集的序列来控制功能相关序列模式。

正文

View PDF HTML (experimental)

Abstract:Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcement learning with sparse autoencoder features (RL-SAE), a post-training method that rewards the generation of sequences that activate specified feature sets. Across eight IDR design tasks, RL-SAE sequences activate, on average, 90% of 30 targeted features, compared to 24% for activation steering. We demonstrate that RL-SAE improves the predicted subcellular localization and transcriptional activity of generated IDRs compared to steering and supervised fine-tuning, and enables features associated with distinct biological functions to be combined within individual sequences. Thus, IDiom and RL-SAE enable interpretable and composable IDR design through explicit control of function-associated sequence features. More broadly, RL-SAE could extend to other protein design settings where interpretable features provide useful design targets. Code is available at this https URL.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02189 [cs.LG]
  (or arXiv:2610.02189v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02189

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jason Liu [view email]
[v1] Thu, 1 Oct 2026 17:59:20 UTC (5,830 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org