arXiv:cs.AI· Saisab Sadhu, Aadit Sengupta, Vinay kumar Sankarapu, Pratinav Seth·· 3 小时前
移除规则后判决依旧:LLM 合规系统的规则敏感性诊断与审计
Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance Systems
AI 导读
一项研究在五个模型、20 个监管与平台政策领域上测试:删除、替换或否定所给规则后,模型判决几乎不变(OCS 与 ICS-delta 均无明显变化)——其中 guard 模型在自定义规则适配的原生分类体系下规则敏感性最低、准确率仅 51%,远低于通用模型的 90-92%,接近随机水平。
正文
Abstract:Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given. We test this directly across five models and 20 regulatory and platform-policy domains: delete, swap, or negate the governing rule while holding the case fixed, and check whether the verdict changes (OCS) or the model's internal representation of compliance shifts at all (ICS-delta). Neither moves much: models' verdicts are often invariant to substantial perturbations of the supplied rule, and the guard model, evaluated here under a custom-rule adaptation of its native taxonomy, is the least rule-sensitive and least accurate of the five, barely above chance (51%, versus 90-92% for general-purpose models). This reflects easy cases more than blanket neglect: on cases where deleting the rule changes a previously correct model prediction, models do track it closely. Neither better prompting nor direct intervention on the model's internal representations closes this gap. Accuracy alone does not establish that a compliance verdict is grounded in the supplied rule.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY) |
| Cite as: | arXiv:2610.12313 [cs.AI] |
| (or arXiv:2610.12313v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12313 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Saisab Sadhu [view email]
[v1]
Thu, 8 Oct 2026 16:59:40 UTC (1,807 KB)
来源:arXiv:cs.AI · arxiv.org