arXiv:cs.LG(机器学习,全量分类)· Vikram Kher, Jane H. Lee, Anay Mehrotra, Manolis Zampetakis·· 5 小时前AI 评分39
样本选择偏差下回归可识别性的紧致刻画
Learning-Enabled Estimation: Tight Characterizations under Sample Selection Biases
AI 导读
研究给出样本选择偏差下回归何时可识别的完整刻画,明确了选择过程函数形式所需的最小假设。其推论显示,存在选择过滤器本身无法识别、但回归函数仍可识别的场景,超越了"先估计选择过滤器再校正回归"的既有范式。在加强条件下,作者还给出显式收敛率的有限样本估计保证与 oracle 高效算法,并探讨了入场成本拍卖、劳动力市场等复杂选择机制的应用。
正文
Abstract:When can we learn from biased samples? We study regression when outcomes are observed only after passing through selection filters that depend on both covariates and outcomes themselves, a ubiquitous challenge spanning clinical trials with patient dropout, labor markets with self-selection, and auctions with strategic entry. Ignoring such selection yields systematically biased conclusions with real-world consequences. This challenge has a long history in econometrics and statistics, starting with Heckman's seminal two-stage model and followed by numerous generalizations. While these works provide various sufficient conditions for identification, a complete characterization of when such regression is possible has remained elusive.
In this work, we provide a characterization for when regression is possible in the presence of sample selection bias. Our results establish the minimal assumptions required on the functional forms of selection processes under which regression remains possible, which are particularly relevant in modern settings where selection mechanisms are increasingly complex and opaque. As a corollary of our characterization, we show that there are settings where the regression function can be identified even when the selection filter itself cannot. This observation already goes beyond the ``estimate selection filter, then debias regression'' paradigm that is followed by virtually all existing approaches. Under natural strengthenings of our identification conditions, we also establish finite-sample estimation guarantees with explicit convergence rates and provide oracle-efficient algorithms. This yields the first general-purpose estimation method for this broad class of selection problems. Finally, we explore the implications of our results for several well-studied econometric settings with complex selection mechanisms such as auctions with entry costs and labor markets.
| Comments: | Presented at EC 2026 |
| Subjects: | Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2609.38608 [cs.LG] |
| (or arXiv:2609.38608v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38608 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jane Lee [view email]
[v1]
Tue, 29 Sep 2026 22:10:50 UTC (203 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org