跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Aaron Rose, Carissa Cullen, Sahar Abdelnabi, Philip Torr, Brandon Gary Kaplowitz, Christian Schroeder de Witt·· 18 小时前AI 评分55

NARCBench:通过多智能体可解释性检测 LLM 智能体合谋

Detecting Multi-Agent Collusion Through Multi-Agent Interpretability

AI 导读

研究者提出 NARCBench 基准和五种探测技术,用模型激活的线性探针在群体层面检测多智能体系统中的隐蔽合谋,并在 Qwen3-32B、Llama-3.1-70B、DeepSeek-R1 32B、GPT-OSS-20B 四个开源权重模型上评估。

正文

View PDF HTML (experimental)

Abstract:As LLM agents are increasingly deployed in multi-agent systems, they introduce risks of covert coordination that may evade standard forms of human oversight. While linear probes on model activations have shown promise for detecting deception in single-agent settings, collusion is inherently a multi-agent phenomenon, and the use of internal representations for detecting collusion between agents remains unexplored. We introduce NARCBench, a benchmark for evaluating collusion detection under environment distribution shift, and propose five probing techniques that aggregate per-agent deception scores to classify scenarios at the group level, evaluated across four open-weight models (Qwen3-32B, Llama-3.1-70B, DeepSeek-R1 32B, GPT-OSS-20B) and six probe architectures. We frame this as a distributed anomaly detection problem, identifying three collusion signatures that map onto distinct anomaly types and detection paradigms. Every model reaches 1.00 AUROC in-distribution; on our strongest model (Llama-3.1-70B), our five probing techniques achieve 0.73 to 0.93 AUROC when transferred zero-shot to structurally different multi-agent scenarios and 0.99 to 1.00 on a steganographic blackjack card-counting task, with detection performance scaling with model capability. We find that no single probing technique dominates across all collusion types, consistent with the framework's prediction that different anomaly types require different detection paradigms. This work takes a step toward multi-agent interpretability: extending white-box inspection from single models to multi-agent contexts, where detection requires aggregating signals across agents. These results suggest that model internals provide a complementary signal to text-level monitoring for detecting multi-agent collusion. Code and data available at this https URL.
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
Cite as: arXiv:2604.01151 [cs.AI]
  (or arXiv:2604.01151v3 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2604.01151

arXiv-issued DOI via DataCite

Submission history

From: Brandon Kaplowitz [view email]
[v1] Wed, 1 Apr 2026 17:08:05 UTC (93 KB)
[v2] Sat, 9 May 2026 19:42:28 UTC (833 KB)
[v3] Thu, 1 Oct 2026 17:58:51 UTC (858 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org