跳到正文
arXiv:cs.CL· Masaaki Nakatsu (AO, Inc. / OrbLabs AG), Reno Wang (AO, Inc.)·· 3 小时前

多智能体 LLM 谈判的宪法门控与确定性恢复:针对有状态对抗 Gatekeeper 的消融实验

Constitutional Gating and Deterministic Recovery for Multi-Agent LLM Negotiation: Ablations Against a Stateful Adversarial Gatekeeper

AI 导读

研究团队用 Gemini 2.5 Pro 智能体和 Claude Haiku 4.5 Gatekeeper 做了 30 次运行,测试由 5-Pillar 运行时宪法、4 层 swarm 与 Cognitive Annealing 组成的三段式控制栈。

正文

View PDF HTML (experimental)

Abstract:Multi-agent LLM systems negotiating with a stateful counterpart waste model calls in three ways: polite loops that never meet the counterpart's hidden acceptance condition, malformed outputs that trigger retries, and compliance deadlocks in which the counterpart demands something the agent must refuse. We study a three-part control stack - a 5-Pillar runtime constitution, a 4-tier swarm (Director, three-agent majority vote, Monitor, schema hard gate) and Cognitive Annealing (deterministic deadlock detection, atomic purge of the agent-side context, a canonical recovery message) - against a released adversarial Gatekeeper whose acceptance rules are fixed regular expressions and whose LLM only renders reply text. The testbed has a known solution: it measures whether the stack executes a constitution-aligned strategy against swarm drift and recovers from deadlock, not whether it discovers anything. In five runs per configuration (30 runs; Gemini 2.5 Pro agents, Claude Haiku 4.5 Gatekeeper) we find: (i) the constitution and Director make an acceptable framing possible but not reliable - 0/5 baseline unlocks versus 1/5 and 2/5 with the constitution; when the swarm unlocks it does so in one turn with 7-8 calls and about 15k tokens (67-73% below baseline); when it does not, it costs 17-38% more; (ii) the Monitor and hard gate do not reduce unlocks and leave an audit trail; (iii) under a honeytrap-to-compliance deadlock, LLM-only steering escapes 0 of 5 times while atomic purge plus a canonical strike escapes 5 of 5 (Fisher $p = 0.008$) at the same call budget, with zero calls for the strike. LLM-written strikes failed the deterministic pre-flight 5 of 5 times although an LLM Monitor had approved four. Pre-registered hypotheses on average call and token reduction were not supported. Cost is bounded in every arm by deterministic stop rules; the stack adds recovery at no extra model cost.
Comments: 29 pages, 3 figures. The Gatekeeper, agents, constitution, lexicon, 30 run logs and analysis scripts are released (see Appendix F). Companion paper: arXiv:2610.09772
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11542 [cs.CL]
  (or arXiv:2610.11542v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.11542

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Masa Nakatsu [view email]
[v1] Thu, 8 Oct 2026 09:10:45 UTC (133 KB)

来源:arXiv:cs.CL · arxiv.org