arXiv:cs.LG· Dongxia Wu, Mingyu Li, Yuhui Zhang, Anurendra Kumar, Emma Lundberg, Serena Yeung-Levy, Emily B. Fox·· 5 小时前AI 评分32
PerturbCellRL:通过后训练扰动生成器对齐分布并锚定生物学
PerturbCellRL: Aligning Distributions and Grounding Biology via Post-Training Perturbation Generators
AI 导读
PerturbCellRL 是一个用强化学习对单细胞扰动生成器进行后训练的框架,通过逐细胞奖励将群体层面的分布偏差转化为细胞级反馈。其核心是基因表达能量见证器,并证明该奖励的策略梯度指向更好的分布对齐。在遗传与化学扰动基准上,该方法显著改善了分布对齐并更忠实地恢复了通路富集模式。
正文
Abstract:Single-cell perturbation models can reduce costly wet-lab screening by predicting how cells respond transcriptionally to interventions. Recent advances in flow-matching have enabled population-level prediction of cellular responses. However, flow-matching training can fail to recover certain target distributions even within the model family, limiting its ability to capture cellular heterogeneity. We first prove that post-training can recover these distributions, then introduce PerturbCellRL, a reinforcement learning framework that post-trains single-cell perturbation generators using per-cell rewards. The central component is a gene-expression energy witness that translates population-level discrepancies into per-cell feedback. We further prove that this reward's policy gradient points toward better distributional alignment. Two complementary rewards, calibrated on real cells, penalize atypical expression profiles and insufficient pathway-level responses to perturbations. Across genetic and chemical perturbation benchmarks, PerturbCellRL substantially improves distributional alignment and recovers pathway enrichment patterns more faithfully. These results establish reward-guided post-training as an effective strategy for improving both distributional accuracy and biological fidelity in perturbation prediction.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2606.27752 [cs.LG] |
| (or arXiv:2606.27752v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2606.27752 arXiv-issued DOI via DataCite |
Submission history
From: Dongxia Wu [view email]
[v1]
Fri, 26 Jun 2026 06:15:04 UTC (4,621 KB)
[v2]
Fri, 2 Oct 2026 15:36:22 UTC (2,970 KB)
来源:arXiv:cs.LG · arxiv.org