arXiv:cs.LG(机器学习,全量分类)· Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik·· 14 小时前AI 评分51
研究检验生物推理模型何时真正使用其生物输入
When Do Biological Reasoning Models Use Their Biological Inputs?
AI 导读
arXiv 论文(arXiv:2610.00898)测试六个生物推理模型在 DNA、蛋白质和单细胞任务上是否真正利用生物输入。
正文
Abstract:Biological reasoning models use post-training to connect LLMs to biological foundation model representations and biological text. Their benchmark accuracy is taken as evidence that LLMs reason over these inputs. We test this assumption in six biological reasoning models across DNA, protein, and single-cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00898 [cs.LG] |
| (or arXiv:2610.00898v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00898 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ada Fang [view email]
[v1]
Thu, 1 Oct 2026 01:27:31 UTC (1,863 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org