跳到正文
arXiv:cs.CL· Pavan Maddula·· 3 小时前AI 评分38

五款开放权重大模型在非规范输入下的四态安全评估

Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs

AI 导读

研究者发布 ASRD 数据集,含 2100 条提示词、覆盖七类表面形式,对五个开放权重语言模型生成 10500 条回复,并用四态评分标准分类。

正文

View PDF HTML (experimental)

Abstract:Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families. Five open-weight language models are evaluated across these prompts, producing 10,500 responses. The Quad-State Evaluation Rubric classifies each response into one of four outcomes: harmful compliance, safe response, comprehension failure, or indeterminate. Emoji and invisible Unicode variations cause almost no comprehension failure, with pooled harmful compliance of 20.27% and 17.20% against a 22.87% baseline that is driven mainly by Mistral 7B, whereas leetspeak, encoded wrappers, and hybrid transformations score 2.40%, 0.13%, and 2.40% while comprehension failure rises to 36.47%, 65.60%, and 34.47%. Inspection of raw model outputs reveals three response behaviors: hallucinated benignity, structural collapse, and language drift. Project page: this http URL
Comments: Accepted at the NeurIPS 2026 Workshop on Self-Evolving Diversity-Driven Search for Robust AI Systems (EvoRobust). 12 pages, 11 tables. Project page: this https URL Dataset: this https URL
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
Cite as: arXiv:2610.09033 [cs.CL]
  (or arXiv:2610.09033v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.09033

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Pavan Maddula [view email]
[v1] Tue, 6 Oct 2026 19:33:24 UTC (20 KB)

来源:arXiv:cs.CL · arxiv.org