跳到正文
arXiv:cs.LG· Dongxia Wu, Mingyu Li, Yuhui Zhang, Anurendra Kumar, Emma Lundberg, Serena Yeung-Levy, Emily B. Fox·· 5 小时前AI 评分32

PerturbCellRL:通过后训练扰动生成器对齐分布并锚定生物学

PerturbCellRL: Aligning Distributions and Grounding Biology via Post-Training Perturbation Generators

AI 导读

PerturbCellRL 是一个用强化学习对单细胞扰动生成器进行后训练的框架,通过逐细胞奖励将群体层面的分布偏差转化为细胞级反馈。其核心是基因表达能量见证器,并证明该奖励的策略梯度指向更好的分布对齐。在遗传与化学扰动基准上,该方法显著改善了分布对齐并更忠实地恢复了通路富集模式。

正文

View PDF HTML (experimental)

Abstract:Single-cell perturbation models can reduce costly wet-lab screening by predicting how cells respond transcriptionally to interventions. Recent advances in flow-matching have enabled population-level prediction of cellular responses. However, flow-matching training can fail to recover certain target distributions even within the model family, limiting its ability to capture cellular heterogeneity. We first prove that post-training can recover these distributions, then introduce PerturbCellRL, a reinforcement learning framework that post-trains single-cell perturbation generators using per-cell rewards. The central component is a gene-expression energy witness that translates population-level discrepancies into per-cell feedback. We further prove that this reward's policy gradient points toward better distributional alignment. Two complementary rewards, calibrated on real cells, penalize atypical expression profiles and insufficient pathway-level responses to perturbations. Across genetic and chemical perturbation benchmarks, PerturbCellRL substantially improves distributional alignment and recovers pathway enrichment patterns more faithfully. These results establish reward-guided post-training as an effective strategy for improving both distributional accuracy and biological fidelity in perturbation prediction.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2606.27752 [cs.LG]
  (or arXiv:2606.27752v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2606.27752

arXiv-issued DOI via DataCite

Submission history

From: Dongxia Wu [view email]
[v1] Fri, 26 Jun 2026 06:15:04 UTC (4,621 KB)
[v2] Fri, 2 Oct 2026 15:36:22 UTC (2,970 KB)

来源:arXiv:cs.LG · arxiv.org