跳到正文
arXiv:cs.CL· Qiaoyuan Zheng, Yiqu Yang·· 4 小时前

近乎持平的 LLM 排名对 Family-DIF 引导的基准重组是否稳健?

Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

AI 导读

一项基于五个基准题级响应的研究发现,跨模型家族差距不足一个百分点时,30.9%–47.1% 的模型对会在基准重组后排名反转,远超随机子采样的中位数 16.9–28.6 个百分点(p=.001),仅第五个基准未见显著超额(-0.9 个百分点,p=.689)。

正文

View PDF HTML (experimental)

Abstract:Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test this using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional item-response theory (MIRT). In owner-disjoint folds, one owner half identifies items with low residual differential item functioning across model families (low-DIF). These items are used to score models in the other half with frozen weights that preserve the benchmark's composition across metadata-defined item groups and easiness strata. Equally short matched-random subtests provide a baseline for variation due to item subsampling. Full-benchmark and low-DIF rankings remain strongly correlated ($\tau_b=.900$--$.948$). Yet in four of five benchmarks, 30.9--47.1\% of cross-family pairs initially within one percentage point reverse order, exceeding their matched-random medians by 16.9--28.6 percentage points (all $p=.001$). The fifth benchmark shows no reliable excess ($-0.9$ points, $p=.689$). The pattern survives all pre-specified population perturbations, and residual item--family signatures replicate across owner halves; however, no family shows a consistent advantage across benchmarks. Thus, globally stable rankings can still leave individual near-tie orderings sensitive to benchmark composition, and sub-one-point leaderboard gaps should be accompanied by evidence that the implied ordering is composition-robust.
Comments: Accepted to NeurIPS 2026 TAE Workshop; Website:this https URL
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2609.00482 [cs.CL]
  (or arXiv:2609.00482v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.00482

arXiv-issued DOI via DataCite

Submission history

From: Yiqu Yang [view email]
[v1] Mon, 31 Aug 2026 23:29:50 UTC (168 KB)
[v2] Wed, 7 Oct 2026 23:17:37 UTC (354 KB)

来源:arXiv:cs.CL · arxiv.org