跳到正文
arXiv:cs.AI· Chenglin Yang·· 6 小时前AI 评分59

arXiv 论文实测:Agent 动作的确定性规则与 LLM 审查层并非独立失效

Evaluate the Stack, Not the Layer: Do Deterministic and LLM Gates for Agent Actions Fail Independently?

AI 导读

arXiv 论文(arXiv:2610.07359)在来自三个语料库的 1119 条标注 Agent 动作上,检验运行时门控层错误相乘的假设。结果显示任意两个 LLM 审查层组合仅相当于约 1.2 到 1.4 个乘法等效层,规则层加一个审查层为 1.86 到 2.09 层;一个云端规则包将规则层单独 miss 率降低 20% 却未带来新的联合覆盖。

正文

View PDF HTML (experimental)

Abstract:Runtime gates for agent tool calls are stacked on the assumption that their errors multiply. We test it on 1,119 labelled agent actions from three corpora, without an adaptive adversary. The stack has one deterministic rule layer and four LLM judges, three of them re-collected with the served model recorded on every call. We read each stack as a number of multiplication-equivalent layers, n_mult, with its floor under perfect coupling. Under the STRICT miss definition (escalation to a human scored as not stopped), any two judges compose to about 1.2 to 1.4 layers ({\phi} median +0.430, 6 of 6 pairs significant, floors 1.02 to 1.17). The rule layer plus one judge composes to 1.86 to 2.09 layers ({\phi} median +0.014, 0 of 4 significant, floors 1.01 to 1.09). Under PRIMARY (escalation scored as caught) the bands are 1.21 to 1.57 and 1.80 to 2.13. Intervals separate on the pooled data, point estimates split on each corpus, and a third-vendor judge lands in the judge band. Solo accuracy does not predict what a layer adds: a cloud rule pack lowers the rule layer's solo miss rate by 20% and adds no new joint coverage. The difficulty share of judge coupling is not identifiable: 31.8% to 61.8% depending on the probe and the miss definition. One judge tier was served by an unrequested model version in 50 of 112 batches, concentrated on the external corpus. That event overturned a pre-declared analysis rule, and the scoring of review verdicts reversed five conclusions. We report both.
Comments: 15 pages, 1 figure, 10 tables. Artifact (data, scripts, provenance): this https URL
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07359 [cs.AI]
  (or arXiv:2610.07359v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07359

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Chenglin Yang [view email]
[v1] Mon, 5 Oct 2026 20:26:58 UTC (61 KB)

来源:arXiv:cs.AI · arxiv.org