跳到正文
arXiv:cs.AI· Min-Young Yu, Tony Kim, Jang Won Choi·· 3 小时前

NOMOS:将书面策略编译为 LLM 智能体静态验证工具调用门禁

NOMOS: Compiling Written Policies into Statically Verified Tool-Call Gates for LLM Agents

AI 导读

NOMOS 是一个四遍编译器,可将自然语言策略转为确定性的工具调用门禁,仅靠工具 schema 级静态检查即可修复或拒绝 37%(航空)和 13%(零售)的候选规则,无需证明器、求解器或 LLM。

正文

View PDF HTML (experimental)

Abstract:Tool-using LLM agents violate the policies they are deployed to enforce, often silently. Prior defenses hand-write rules, query an LLM verifier per action, or compile policies through heavyweight formal machinery. Naive compilation fails: extracted rules block the tool satisfying their own precondition, or read arguments their tool lacks. NOMOS, a four-pass compiler, turns a natural-language policy into a deterministic tool-call gate; static verification with tool-schema-level checks alone (no prover, solver, or LLM) repairs or rejects 37% (airline) and 13% (retail) of candidates, without which most shipped rules are inoperable. Replaying compiled rules over undefended transcripts flags bindings that refuse legitimate work (a development binding refused 95.9% of task-passing calls); no evaluation binding is flagged. On $\tau^2$-bench the gate cuts violations of reference-encoded clauses among state-changing calls from 66.3% to 2.6% (airline) and 30.8% to 6.9% (retail), raising airline task success significantly for $2 \le k \le 4$; a 26B on-premise compilation is not significantly worse than hand-written or frontier-compiled rules. Unlike AgentDojo's shipped defenses, it reaches a zero attack success rate (ASR) on banking, where nine attack families collapse onto three structural rules. On the other three suites its ASR is at most 3.6%, from goals with no tool call to govern and one write admitted by a binding weaker than its clause; a second agent model, Llama-3.3-70B, reproduces the effect on both benchmarks. Decisions take microseconds without an LLM call, at a domain-dependent benign-utility cost; compilation runs on-premise on open-weight gemma-4-26B.
Comments: 28 pages, 5 figures, 19 tables. Submitted to IEEE Access
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.11030 [cs.CR]
  (or arXiv:2610.11030v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.11030

arXiv-issued DOI via DataCite (pending registration)

Related DOI: https://doi.org/10.5281/zenodo.22123420

DOI(s) linking to related resources

Submission history

From: Min-Young Yu [view email]
[v1] Thu, 8 Oct 2026 00:23:22 UTC (382 KB)

来源:arXiv:cs.AI · arxiv.org