arXiv:cs.LG· Hazar Yueksel·· 4 小时前AI 评分31
损失差分条件互信息的精度—信息权衡
An Accuracy--Information Tradeoff for Loss-Difference Conditional Mutual Information
AI 导读
研究证明,在维度至少与 n 线性相关的缩放符号立方体上的乘积分布中,对零处斜率非零的平滑凸损失(如 logistic loss)加幂次 r≥2 的正则项,最优样本量 n≍ε^{-2+2/r} 下期望超额风险不超过 ε 的任意 proper learner,其最坏情况 ld-CMI 为 n bits 量级。
正文
Abstract:Loss-difference conditional mutual information (ld-CMI) uses the smallest of the standard observations in the supersample hierarchy of generalization bounds: it measures what a learner's loss differences reveal about which candidate of each pair it was trained on. Accuracy is known to force information into the model; data processing does not carry such lower bounds to losses. We show, by bounding three moments of the loss differences, that accuracy also forces ld-CMI. For linear predictors with a smooth convex loss of nonzero slope at zero, such as the logistic loss, plus a regularizer whose curvature and growth are both of power $r\ge2$, on product distributions over a scaled sign cube in dimension at least linear in $n$, every proper learner with expected excess risk at most $\varepsilon$ on these distributions at the optimal sample size $n\asymp\varepsilon^{-2+2/r}$ has worst-case ld-CMI of order $n$ bits, and $\Theta(n/(1+(\tau/\varepsilon)^2))$ bits under Gaussian noise of standard deviation $\tau$ on the loss differences. The same holds without a regularizer, at $n\asymp\varepsilon^{-2}$. Consequently, range-scaled ld-CMI bounds cannot vanish on these distributions, although every proper learner's generalization gap is $O(n^{-1/2})$. We also show that model-level information does not determine noisy loss-difference information, and that the growth, slope and dimension conditions are needed, the last up to a logarithm.
| Comments: | 61 pages, of which 8 pages main text. Code, data and the Lean 4 formalization are in the ancillary files |
| Subjects: | Machine Learning (cs.LG); Information Theory (cs.IT); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.09206 [cs.LG] |
| (or arXiv:2610.09206v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09206 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hazar Yueksel [view email]
[v1]
Tue, 6 Oct 2026 23:07:16 UTC (3,308 KB)
来源:arXiv:cs.LG · arxiv.org