跳到正文
arXiv:cs.AI· Andrea Wynn (Nathan), Harsh Satija (Nathan), Seokhyun (Nathan), Baek, Anqi Liu, Eric Nalisnick, Gillian K. Hadfield·· 3 小时前

LLM 规范能力的基础、设计与挑战:多智能体社区辩论中的规范推断研究

Reading the Room: Foundations, Design, and Challenges of Normative Competence in LLMs

AI 导读

研究者提出多智能体社区辩论设置,用合成规范控制辩论准入,以独立于预训练暴露来评估 LLM 的规范能力。基线 LLM 智能体即使学会规范能提升准确率也仍然失败;引入规范推断模块后,规范遵循对规范风格和模块所依托的模型高度敏感,且当真实规范伴随非规范行为时,智能体会不加区分地模仿这些噪声,即使明确惩罚也无济于事。

正文

View PDF HTML (experimental)

Abstract:Human communities are governed by normative systems: shared standards that produce \textit{norms} dictating acceptable behavior, enforced through community sanctioning. Aligning increasingly autonomous AI systems with these norms is a central alignment challenge, complicated by the fact that norms are vast in number, change quickly, and are often arbitrary (e.g., dress or language conventions). Thus, alignment requires \textit{normative competence}: the ability to discern from interaction alone what norms a community enforces without relying on static pretrained knowledge. We introduce a multi-agent community debate setting, where access to debate is governed by synthetic norms, to study normative competence in isolation from pretraining exposure. We show that baseline LLM agents fail to learn norms even when doing so would improve their accuracy. We then experiment with various \textit{normative modules} -- architectural components for norm inference -- finding that norm-following is highly sensitive to both the style of norm and the model powering the normative module, suggesting a lack of generalizability. Furthermore, when idiosyncratic, non-normative behaviors accompany the true norm, LLM agents exhibit an unselective attribution failure: they indiscriminately copy idiosyncratic noise alongside enforced rules, a pattern that persists even when imitating unnecessary behaviors is explicitly penalized. To the best of our knowledge, our work is the first to operationalize and evaluate normative competence in LLMs, demonstrating that current AI systems excel at behavioral mimicry but lack the capacity to discern socially enforced order.
Subjects: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
Cite as: arXiv:2610.10906 [cs.AI]
  (or arXiv:2610.10906v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.10906

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Andrea Wynn [view email]
[v1] Wed, 7 Oct 2026 21:02:52 UTC (404 KB)

来源:arXiv:cs.AI · arxiv.org