跳到正文
arXiv:cs.AI· Saisab Sadhu, Aadit Sengupta, Vinay kumar Sankarapu, Pratinav Seth·· 3 小时前

移除规则后判决依旧:LLM 合规系统的规则敏感性诊断与审计

Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance Systems

AI 导读

一项研究在五个模型、20 个监管与平台政策领域上测试:删除、替换或否定所给规则后,模型判决几乎不变(OCS 与 ICS-delta 均无明显变化)——其中 guard 模型在自定义规则适配的原生分类体系下规则敏感性最低、准确率仅 51%,远低于通用模型的 90-92%,接近随机水平。

正文

View PDF HTML (experimental)

Abstract:Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given. We test this directly across five models and 20 regulatory and platform-policy domains: delete, swap, or negate the governing rule while holding the case fixed, and check whether the verdict changes (OCS) or the model's internal representation of compliance shifts at all (ICS-delta). Neither moves much: models' verdicts are often invariant to substantial perturbations of the supplied rule, and the guard model, evaluated here under a custom-rule adaptation of its native taxonomy, is the least rule-sensitive and least accurate of the five, barely above chance (51%, versus 90-92% for general-purpose models). This reflects easy cases more than blanket neglect: on cases where deleting the rule changes a previously correct model prediction, models do track it closely. Neither better prompting nor direct intervention on the model's internal representations closes this gap. Accuracy alone does not establish that a compliance verdict is grounded in the supplied rule.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)
Cite as: arXiv:2610.12313 [cs.AI]
  (or arXiv:2610.12313v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.12313

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Saisab Sadhu [view email]
[v1] Thu, 8 Oct 2026 16:59:40 UTC (1,807 KB)

来源:arXiv:cs.AI · arxiv.org