arXiv:cs.LG· Jyotishka Ray Choudhury, Kabir Aladin Verchand, Richard J. Samworth, Ashwin Pananjady·· 4 小时前AI 评分29
缺失协变量的弱假设逻辑回归:基于 Z-estimation 与单调算子的随机逼近方法
Assumption-lean logistic regression with missing covariates
AI 导读
针对协变量分布未知且存在缺失的逻辑回归参数估计问题,研究者提出一种基于 Z-estimation 与新型单调算子的随机逼近方法,在协变量完全随机缺失假设下实现参数速率下的信号恢复,且计算高效。理论刻画了估计量的 ℓ₂² 风险与缺失模式的精细关系,并给出信息论下界证明该依赖在极小极大意义下是本质的。该方法始终优于忽略含缺失数据的“仅完整样本”估计器。
正文
Abstract:Missing covariates are frequently encountered in supervised learning problems, and classical methods for estimation using such data use carefully chosen imputation schemes for missing data, or likelihood approximations that lead to nonconvex $M$-estimation problems. These methods and their relatives are suitable for scenarios in which the covariate distribution is known, and more broadly, have enjoyed tremendous success in linear models. But even in basic nonlinear problems such as logistic regression in moderate dimensions, such methods can experience drastic failure modes when the covariate distribution is unknown.
Motivated by the need for reliable alternatives, we consider the problem of parameter estimation in logistic regression with missing covariates. Crucially, we operate in the assumption-lean setting where the covariate distribution is unknown (but bounded). We design a stochastic approximation method that is based on $Z$-estimation with a novel monotone operator, and establish that our algorithm is computationally efficient and achieves provable signal recovery at parametric rates under the hypothesis that covariates are missing completely at random. Our theory sharply characterizes the $\ell_2^2$ risk of the estimator in terms of the missingness profile, accommodating heterogeneous observation probabilities. Importantly, it shows that our method always outperforms the de facto ``complete-case'' estimator that ignores observations with any missing data. Even in the setting with homogeneous missingness (in which each covariate is observed independently with probability $q$), our bounds exhibit intricate and nonstandard dependence on $q$ that can yield significant improvements over using only complete cases. We complement our upper bounds with new information-theoretic lower bounds that show that this intricate dependence on $q$ is fundamental in a minimax sense.
| Subjects: | Machine Learning (stat.ML); Information Theory (cs.IT); Machine Learning (cs.LG); Statistics Theory (math.ST); Methodology (stat.ME) |
| MSC classes: | 62J12 (Primary), 62C20, 62L20 (Secondary) |
| Cite as: | arXiv:2610.07292 [stat.ML] |
| (or arXiv:2610.07292v1 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07292 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jyotishka Ray Choudhury [view email]
[v1]
Mon, 5 Oct 2026 19:28:57 UTC (90 KB)
来源:arXiv:cs.LG · arxiv.org