arXiv:cs.AI· Christopher M. Stewart, Preston Botter, Natalie Sarabosing, Muye Zhang, Rachel Phinnemore, Shalini Ghosh, Hong Shen, Hoda Heidari·· 3 小时前
研究质疑 HELM Safety 能否测量单一"有害拒绝"属性
Searching for "Harmful Refusal": A Psychometric Audit of an AI Safety Benchmark
AI 导读
一项心理测量学审计发现,HELM Safety 中可能对应"有害拒绝"的四个数据集中有三个已饱和,剩余数据集 HarmBench 经多维项目反应理论建模后,强烈表明其并未测量单一属性。差分项目功能分析还发现,来自不同开发者的模型在拒绝能力相同时,部分题目得分仍存在差异,这种模式与聚合效应一致但不足以排除开发者间的领域差异。研究者指出,任何对数据集和题目求平均的安全评分都可能掩盖饱和并混淆不同行为。
正文
Abstract:Safety benchmarks typically report one overall score for a suite of datasets, each of which may target one or more safety-related attributes, so models with similar overall scores can have very different attribute profiles. Comparing models is more tractable at the level of individual attributes, yet it is often unclear whether even a single dataset's scores isolate any single attribute. One plausible candidate for such an attribute is harmful refusal, a model's tendency to refuse dangerous or policy-violating prompts. We examine whether it constitutes a single, measurable attribute in HELM Safety. Using a construct validity framework that stipulates that an attribute must exist before a test can measure it, we start with HELM Safety's four datasets that might plausibly target harmful refusal, but find that three are saturated. We subject the remaining dataset, HarmBench, to two psychometric tests to determine if a single attribute like harmful refusal could stand behind its score. First, multidimensional item response theory modeling strongly suggests that HarmBench does not measure a singular attribute. Second, a differential item functioning analysis finds items where models from different developers with the same refusal ability score differently. These flags largely disappear under scope-specific matching, a pattern consistent with aggregation effects but not sufficient to rule out domain-specific developer differences. Zooming out, HarmBench collapses distinct harm behaviors into one score, and the overall HELM safety aggregate further collapses HarmBench and scores from other datasets into a single top-line number. Any safety score that averages over datasets and items can hide saturation and conflate behaviors this way. We argue that a score should earn its single-attribute reading before models are compared with it.
| Comments: | COLM 2026, AI for Measurement Science (AIMS) Workshop |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.12409 [cs.AI] |
| (or arXiv:2610.12409v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12409 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Natalie Sarabosing [view email]
[v1]
Thu, 8 Oct 2026 17:46:43 UTC (123 KB)
来源:arXiv:cs.AI · arxiv.org