arXiv:cs.LG· JM Gorriz·· 5 小时前AI 评分36
准确率何时构成证据?泛化、验证与信息融合的统一理论
When Is Accuracy Evidence? A Unified Theory of Generalisation, Validation, and Information Fusion
AI 导读
研究者提出 gamma-CUBV 框架,用累积量包络 gamma(lambda) 控制泛化差距的矩生成函数,统一覆盖 Hoeffding、Bernstein、依赖感知、PAC-Bayesian 与异构多源融合等风险界。
正文
Abstract:K-fold cross-validation (CV) is widely used as evidence of out-of-sample performance, although folds are neither independent experiments nor equally informative under heterogeneous data. Cross Upper-Bound Validation (CUBV) replaces point-wise CV accuracy by conservative upper bounds on true risk. Here we generalise CUBV through a single exponential framework in which the moment-generating function of the generalisation gap is controlled by a cumulant envelope gamma(lambda). This yields a family of risk bounds covering Hoeffding-, Bernstein-, dependency-aware, PAC-Bayesian, and heterogeneous source-fusion settings. For K-fold CV, dependence between fold-wise gaps is modelled through a joint sub-Gaussian proxy matrix. Under equicorrelation, this gives an effective number of folds, Keff = K/[1+(K-1)rho], showing that increasing K does not necessarily increase statistical evidence when folds are strongly dependent. The framework is also extended to posterior distributions over predictors and weighted multi-source fusion, where weights are selected by minimising an upper bound on future risk rather than empirical error alone. Experiments with trained linear classifiers on heterogeneous multimodal Gaussian mixtures compare K-fold CV with full-sample resubstitution plus risk correction. Bounds are evaluated by coverage and tightness. In low-dimensional small-sample settings, K-fold partitioning can increase uncertainty because individual folds under-represent minority modes, while corrected resubstitution can remain valid and tighter; this effect disappears as sample size increases. Overall, gamma-CUBV separates observed performance, uncertainty, dependence, model complexity, and confidence into explicit terms, providing a unified route from CV scores to risk statements and a principled validation criterion for heterogeneous small-sample applications such as neuroimaging.
| Comments: | 52 pages, 30 figures |
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG); Data Analysis, Statistics and Probability (physics.data-an) |
| Cite as: | arXiv:2610.03465 [stat.ML] |
| (or arXiv:2610.03465v1 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03465 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Juan Manuel Gorriz Saez [view email]
[v1]
Fri, 2 Oct 2026 15:40:28 UTC (9,497 KB)
来源:arXiv:cs.LG · arxiv.org