跳到正文
arXiv:cs.LG· Huali Zhao, Molei Liu, Tianying Wang·· 4 小时前AI 评分29

Data Fusion for Errors-in-Variables:基于条件可迁移性的数据融合估计

Data Fusion for Errors-in-Variables

AI 导读

研究针对目标研究仅含单个易错替代变量、外部源研究提供重复替代测量的 errors-in-variables 问题,提出条件可迁移性假设与复制误差条件,识别目标条件测量误差分布。

正文

View PDF HTML (experimental)

Abstract:We study errors-in-variables problems in which a target study contains only a single error-prone surrogate of an unobserved exposure, while an external source study provides repeated surrogate measurements from a different population. The measurement error distribution is allowed to depend on the observed error-free variables, and the error-free variable distribution itself may differ between studies. We introduce a conditional transportability assumption that enables the use of external repeated measurements under source-target heterogeneity. Together with additional replicate-error conditions, it identifies the target conditional measurement-error distribution. Building on this identification result, we develop a data-fusion estimator for a broad class of target functionals. The estimator combines conditional deconvolution, flexible nuisance estimation, and orthogonal correction that reduces first-order sensitivity to nuisance estimation. For the proposed estimator, we develop a unified spectral theory covering both diffuse-spectrum and finite atomic-spectrum target functionals, derive a general asymptotic expansion, and establish consistency and target-specific convergence-rate bounds. The resulting convergence-rate bounds depend jointly on the spectral properties of the measurement error, the latent exposure, and the target functional. For finite atomic-spectrum targets, we further establish joint Gaussian and bootstrap limits, yielding inference for smooth moment transformations under an additional centering condition. In the reported simulations, Fuse-EIV has small bias for the primary exposure-related coefficient. Applications to the National Health and Nutrition Examination Survey illustrate how accounting for population heterogeneity and error heteroscedasticity can change empirical conclusions.
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST); Methodology (stat.ME)
Cite as: arXiv:2610.07048 [stat.ML]
  (or arXiv:2610.07048v1 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2610.07048

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tianying Wang [view email]
[v1] Mon, 5 Oct 2026 02:22:43 UTC (119 KB)

来源:arXiv:cs.LG · arxiv.org