跳到正文
arXiv:cs.LG· JM Gorriz·· 5 小时前AI 评分36

准确率何时构成证据?泛化、验证与信息融合的统一理论

When Is Accuracy Evidence? A Unified Theory of Generalisation, Validation, and Information Fusion

AI 导读

研究者提出 gamma-CUBV 框架,用累积量包络 gamma(lambda) 控制泛化差距的矩生成函数,统一覆盖 Hoeffding、Bernstein、依赖感知、PAC-Bayesian 与异构多源融合等风险界。

正文

View PDF HTML (experimental)

Abstract:K-fold cross-validation (CV) is widely used as evidence of out-of-sample performance, although folds are neither independent experiments nor equally informative under heterogeneous data. Cross Upper-Bound Validation (CUBV) replaces point-wise CV accuracy by conservative upper bounds on true risk. Here we generalise CUBV through a single exponential framework in which the moment-generating function of the generalisation gap is controlled by a cumulant envelope gamma(lambda). This yields a family of risk bounds covering Hoeffding-, Bernstein-, dependency-aware, PAC-Bayesian, and heterogeneous source-fusion settings. For K-fold CV, dependence between fold-wise gaps is modelled through a joint sub-Gaussian proxy matrix. Under equicorrelation, this gives an effective number of folds, Keff = K/[1+(K-1)rho], showing that increasing K does not necessarily increase statistical evidence when folds are strongly dependent. The framework is also extended to posterior distributions over predictors and weighted multi-source fusion, where weights are selected by minimising an upper bound on future risk rather than empirical error alone. Experiments with trained linear classifiers on heterogeneous multimodal Gaussian mixtures compare K-fold CV with full-sample resubstitution plus risk correction. Bounds are evaluated by coverage and tightness. In low-dimensional small-sample settings, K-fold partitioning can increase uncertainty because individual folds under-represent minority modes, while corrected resubstitution can remain valid and tighter; this effect disappears as sample size increases. Overall, gamma-CUBV separates observed performance, uncertainty, dependence, model complexity, and confidence into explicit terms, providing a unified route from CV scores to risk statements and a principled validation criterion for heterogeneous small-sample applications such as neuroimaging.
Comments: 52 pages, 30 figures
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Data Analysis, Statistics and Probability (physics.data-an)
Cite as: arXiv:2610.03465 [stat.ML]
  (or arXiv:2610.03465v1 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2610.03465

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Juan Manuel Gorriz Saez [view email]
[v1] Fri, 2 Oct 2026 15:40:28 UTC (9,497 KB)

来源:arXiv:cs.LG · arxiv.org