arXiv:cs.AI· Haoxiang Wang, Da Yu, Huishuai Zhang·· 4 小时前
超越固定基准与最坏情况攻击:语言模型的动态边界评估
Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models
AI 导读
研究者提出动态边界评估(DBE),主动定位每个 LLM 在随机采样解码下通过概率接近 0.5 的边界,并将其映射到全局可比的难度标尺上。DBE 包含经 9 个参考 LLM 验证的校准题库、仅需 API 级查询的技能引导边界搜索(SGBS)算法,以及统一能力评估协议。该方法覆盖安全(有害请求拒答、过度拒答)、能力(约束指令遵循)与真实性(多轮谄媚抵抗)四类任务,在避免饱和的同时兼容现有数据集。
正文
Abstract:Evaluating large language models (LLMs) today rests on fixed benchmarks that apply the same set of items to any model, producing ceiling and floor effects that mask capability gaps. We argue that the most informative evaluation signal lies at the boundary, where the per-prompt pass probability is near $0.5$ under random-sampling decoding strategy, and propose Dynamic Boundary Evaluation (DBE), which actively locates each model's boundary and places it on a globally comparable difficulty scale. DBE delivers three artifacts: (i) a calibrated item bank covering safety, capability, and truthfulness, with per-item difficulty labels validated across $9$ reference LLMs; (ii) Skill-Guided Boundary Search (SGBS), a search algorithm that finds boundary items for a given target LLM using only API-level query access; and (iii) an evaluation protocol that places a new LLM on a unified ability scale and grows the evaluation set adaptively when the target falls outside the bank's coverage. We instantiate DBE on four categories spanning safety (harmful request refusal, over-refusal), capability (constrained instruction following), and truthfulness (multi-turn sycophancy resistance). The resulting evaluation covers a broader model spectrum without saturation while remaining compatible with existing datasets.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2605.06213 [cs.AI] |
| (or arXiv:2605.06213v3 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2605.06213 arXiv-issued DOI via DataCite |
Submission history
From: Haoxiang Wang [view email]
[v1]
Thu, 7 May 2026 13:15:31 UTC (86 KB)
[v2]
Tue, 26 May 2026 15:14:12 UTC (1 KB) (withdrawn)
[v3]
Thu, 8 Oct 2026 10:38:22 UTC (89 KB)
来源:arXiv:cs.AI · arxiv.org