arXiv:cs.AI· Min-Young Yu, Tony Kim, Jang Won Choi·· 3 小时前
NOMOS:将书面策略编译为 LLM 智能体静态验证工具调用门禁
NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents
AI 导读
NOMOS 是一个四遍编译器,可将自然语言策略转为确定性的工具调用门禁,仅靠工具 schema 级静态检查即可修复或拒绝 37%(航空)和 13%(零售)的候选规则,无需证明器、求解器或 LLM。
正文
Abstract:Tool-using LLM agents violate the policies they are deployed to enforce, often silently. Prior defenses hand-write rules, query an LLM verifier per action, or compile policies through heavyweight formal machinery. Naive compilation fails: extracted rules block the tool satisfying their own precondition, or read arguments their tool lacks. NOMOS, a four-pass compiler, turns a natural-language policy into a deterministic tool-call gate; static verification with tool-schema-level checks alone (no prover, solver, or LLM) repairs or rejects 37% (airline) and 13% (retail) of candidates, without which most shipped rules are inoperable. Replaying compiled rules over undefended transcripts flags bindings that refuse legitimate work (a development binding refused 95.9% of task-passing calls); no evaluation binding is flagged. On $\tau^2$-bench the gate cuts violations of reference-encoded clauses among state-changing calls from 66.3% to 2.6% (airline) and 30.8% to 6.9% (retail), raising airline task success significantly for $2 \le k \le 4$; a 26B on-premise compilation is not significantly worse than hand-written or frontier-compiled rules. Unlike AgentDojo's shipped defenses, it reaches a zero attack success rate (ASR) on banking, where nine attack families collapse onto three structural rules. On the other three suites its ASR is at most 3.6%, from goals with no tool call to govern and one write admitted by a binding weaker than its clause; a second agent model, Llama-3.3-70B, reproduces the effect on both benchmarks. Decisions take microseconds without an LLM call, at a domain-dependent benign-utility cost; compilation runs on-premise on open-weight gemma-4-26B.
| Comments: | 28 pages, 5 figures, 19 tables. Submitted to IEEE Access |
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.11030 [cs.CR] |
| (or arXiv:2610.11030v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11030 arXiv-issued DOI via DataCite (pending registration) |
|
| Related DOI: | https://doi.org/10.5281/zenodo.22123420
DOI(s) linking to related resources |
Submission history
From: Min-Young Yu [view email]
[v1]
Thu, 8 Oct 2026 00:23:22 UTC (382 KB)
来源:arXiv:cs.AI · arxiv.org