跳到正文
arXiv:cs.LG· Raphael C Kim, Jingsen Zhu, Ramin Zabih, Michele Santacatterina·· 6 小时前AI 评分32

好的生成器就是好的决策者吗?通过重定向反事实生成实现通用干预的策略学习

Are Good Generators Good Decision-Makers? Policy Learning for General Interventions via Retargeted Counterfactual Generation

AI 导读

针对生成模型在联合高维干预与高维结果决策场景中的三大挑战,研究提出"重定向反事实生成策略学习"方法:先学习双重稳健的不变反事实生成器,再基于其 rollout 进行策略学习,最后将生成器重定向至所学策略的干预并重新学习策略。

正文

View PDF HTML (experimental)

Abstract:Generative models are increasingly used to support decision-making in complex systems, where interventions may be joint and high-dimensional, and outcomes are high-dimensional. However, using generators for these decision-making settings are challenged by three problems. First, they are often trained on noisy logs with limited intervention data. Second, generator learning can be noisy across different environments. Third, generators are not optimized for the decisions they support. We introduce policy learning via retargeted counterfactual generation, which trains a generator for the decisions it supports in three steps. We (1) learn a doubly-robust, invariant counterfactual generator for high-dimensional interventions and outcomes, (2) conduct policy-learning based on its rollouts, and (3) retarget the generator toward the learned policy's interventions and relearn the policy, so the generator is accurate where decisions are made. Theoretically, our generator's excess counterfactual risk has a doubly robust product remainder, and retargeting removes the worst-case density ratio between the logged and learned policies from the regret. We demonstrate the efficacy of our approach across synthetic data, Cell Painting images of SARS-CoV-2-infected cells, and physical video simulations.
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as: arXiv:2606.07399 [stat.ML]
  (or arXiv:2606.07399v3 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2606.07399

arXiv-issued DOI via DataCite

Submission history

From: Raphael Kim [view email]
[v1] Fri, 5 Jun 2026 15:40:59 UTC (17 KB)
[v2] Fri, 21 Aug 2026 23:12:16 UTC (22 KB)
[v3] Wed, 7 Oct 2026 10:20:51 UTC (147 KB)

来源:arXiv:cs.LG · arxiv.org