arXiv:cs.LG(机器学习,全量分类)· Hongyu Cao, Xinyuan Wang, Arun Vignesh Malarkkan, Kunpeng Liu, Haifeng Chen, Yanjie Fu·· 14 小时前AI 评分26
IGDPR:协变量偏移下的不变引导扩散与原型重加权
Rethinking Data Augmentation under Covariate Shift: Invariant-Guided Diffusion and Prototype Reweighting
AI 导读
针对表格数据稀缺且训练与部署间存在协变量偏移的场景,研究者提出 IGDPR 框架,将问题统一为 AWL-CS(协变量偏移下的增强与加权学习)。该方法用不变势引导扩散采样,使合成样本对齐稳定决策边界,并以原型聚类重加权过滤验证噪声。真实数据实验显示其能提升数据质量,增强学习鲁棒性。
正文
Abstract:In many industrial applications, 1) tabular data is scarce and imbalanced and thus requires synthetic expansion; 2) input distributions drift between training and deployment (covariate shift); 3) validation sets often diverge from unseen test environments; or 4) standard generative models simply mimic outdated source distributions. This learning setting limits the stability of standard augmentation and adaptation pipelines. We generalize the task under such setting as the Augmented and Weighted Learning under Covariate Shift problem (AWL-CS). AWL-CS imposes two critical challenges on existing methods: 1) misleading generative guidance where models optimize for source similarity rather than downstream task relevance, and 2) structural instability of distributional density where reweighting mechanisms overfit to noisy validation signals. To tackle these challenges, we propose IGDPR (Invariant-Guided Diffusion with Prototype Reweighting), a unified framework that synergizes stable synthesis and structural adaptation: i) To achieve task-relevant generation, we steer the diffusion sampling process using invariant potentials to ensure synthetic samples align with stable decision boundaries rather than outdated correlations. ii) To ensure stable adaptation, we develop a prototype-based reweighting strategy that assesses sample reliability through structural clusters instead of isolated points, effectively filtering validation noise. Extensive experiments on real data demonstrate our method improves data quality by augmenting the most beneficial data for robust learning.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.00873 [cs.LG] |
| (or arXiv:2610.00873v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00873 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hongyu Cao [view email]
[v1]
Thu, 1 Oct 2026 00:46:17 UTC (4,271 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org