arXiv:cs.AI· Weiying Chen, Junlong Shen, Zhanyuan Guo, Xiaoou Zhou·· 4 小时前AI 评分45
LLM 在《克苏鲁的呼唤》TRPG 中作为裁决者的规则遵循评估
Assessing Rule Adherence of LLM Adjudicators in Call of Cthulhu TRPG
AI 导读
研究提出 CoC-Seduce 多智能体对抗基准,用 GPT-5.4、Claude Sonnet 4.6、Gemini 3.5 Flash 生成 5,376 个样本,覆盖 4 种世界设定与 16 类技能,测试 22 个裁决模型。结果显示新版本与显式推理均不能可靠提升裁决鲁棒性,伪逻辑框架是最有效的修辞注入手法,世界设定的影响有限。
正文
Abstract:As LLMs are increasingly deployed as autonomous adjudicators in games such as Call of Cthulhu (CoC), robust rule adherence becomes critical when user intent conflicts with system rules. However, as these models are trained to be helpful and compliant, they may be vulnerable to a class of manipulations we term Rhetorical Injection, where adversarial users exploit narrative framing techniques such as pseudo-logical reasoning and authoritative coercion to bypass adjudication logic. We present CoC-Seduce, a multi-agent adversarial benchmark built on CoC, a Tabletop Role-Playing Game (TRPG) in which rules are explicit about which risky actions require adjudication, yet interaction remains entirely in natural language. Three LLMs, i.e., GPT-5.4, Claude Sonnet 4.6, Gemini 3.5 Flash, serve as adversarial generators producing 5,376 samples across 4 world settings and 16 skill categories. We then benchmark 22 target adjudicators against this corpus. Evaluation across 22 models reveals that neither newer releases nor explicit reasoning reliably confer adjudication robustness, that Pseudo-Logic framing is the most effective rhetorical style, and that the world setting, including culturally distant ones, has only a modest effect. Project page: this https URL.
| Comments: | corrected errors, added evaluations of new models, and revised the scope of the paper |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.02802 [cs.CL] |
| (or arXiv:2607.02802v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.02802 arXiv-issued DOI via DataCite |
Submission history
From: Weiying Chen [view email]
[v1]
Thu, 2 Jul 2026 22:25:35 UTC (365 KB)
[v2]
Fri, 2 Oct 2026 05:55:32 UTC (371 KB)
来源:arXiv:cs.AI · arxiv.org