arXiv:cs.LG· Lan Tao, Yongxian He, Shirong Xu, Yidong Ouyang, Guang Cheng·· 3 小时前
超越分布保真度:面向合成表格数据的因果惩罚扩散模型
Beyond Distributional Fidelity: Causal-Penalized Diffusion for Synthetic Tabular Data
AI 导读
研究提出一种因果保真度感知训练框架,在生成目标中加入因果差异惩罚项,并以因果惩罚版 TabDDPM 实例化,通过 on-policy score-function 估计器优化。理论结果表明高统计保真度通常不能保证高因果保真度,并给出了因果正则化提升期望因果保真度的条件。实验在多种处理效应模拟和两个基准数据集上验证了该方法在提升因果保真度的同时保持有竞争力的统计保真度。
正文
Abstract:Synthetic tabular generators are commonly optimized for distributional fidelity, but statistical similarity alone does not guarantee preservation of causal effects. In this paper, we study whether causal fidelity can be improved directly within a fully generative tabular model. Causal Fidelity is defined with respect to a target estimand as the discrepancy between inferential distributions obtained from real and synthetic data, and theoretical results show that high statistical fidelity does not generally imply high causal fidelity. We then propose a causal-fidelity-aware training framework which adds a causal discrepancy penalty to the generative objective. The framework is instantiated with a causal-penalized TabDDPM and optimized using an on-policy score-function estimator. We further establish conditions under which causal regularization improves expected causal fidelity. Experiments across diverse treatment-effect simulations and two benchmark datasets evaluate the ability of our method to improve causal fidelity while preserving competitive statistical fidelity.
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.11407 [stat.ML] |
| (or arXiv:2610.11407v1 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11407 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Lan Tao [view email]
[v1]
Thu, 8 Oct 2026 07:36:24 UTC (1,276 KB)
来源:arXiv:cs.LG · arxiv.org