arXiv:cs.LG(机器学习,全量分类)· Adam Noonan·· 15 小时前AI 评分38
聚类下的超限设计效应:阈值估计的有效样本量
The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering
AI 导读
研究证明,当观测以独立分组形式出现时,分组会把阈值比例的大样本方差乘以 1+(m-1)ρ_I(p),其中 m 为组大小,p 为目标比例,ρ_I(p) 衡量组内两成员是否落在阈值同一侧,且该相关性与分数本身的相关性不同、随目标变化。
正文
Abstract:Suppose we want a cutoff that 90% of a population falls below. We estimate it from a sample, and another sample would give a different cutoff and a different fraction below it. We ask how much that fraction varies when observations come in independent groups, such as pupils in classrooms or sentences in news articles. We prove that grouping multiplies its large-sample variance by $1+(m-1)\rho_I(p)$, where $m$ is the group size, $p$ is the target fraction, and $\rho_I(p)$ measures whether two members of a group fall on the same side of the cutoff. That correlation can differ from the correlation between the scores themselves, and it changes with the target. We give a direct proof, a counterexample to using score correlation, and an extension to unequal group sizes.
A dataset therefore does not have one effective sample size. How much information it contains depends on the question you ask. In our document experiment, the same 1,000 rows carried about 217 independent observations' worth of information at the median. At the 95th percentile, they carried about 621. Nothing about the dataset changed. We asked it a different question. The number of rows is a property of the dataset. The effective sample size belongs to the analysis.
| Comments: | 22 pages, 2 figures. Lean proofs and code: this https URL |
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG) |
| MSC classes: | 62G15, 62D05 |
| ACM classes: | G.3 |
| Cite as: | arXiv:2608.21262 [stat.ML] |
| (or arXiv:2608.21262v2 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2608.21262 arXiv-issued DOI via DataCite |
Submission history
From: Adam Noonan [view email]
[v1]
Fri, 21 Aug 2026 16:19:50 UTC (150 KB)
[v2]
Thu, 1 Oct 2026 13:22:23 UTC (70 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org