arXiv:cs.LG· Yuhwan Jeong, Jinnyeong Yang, Kuk-Jin Yoon·· 4 小时前AI 评分33
LLM 会依照已知信息行动吗?从伙伴表征到合作行为
Do LLMs Act on What They Know? From Partner Representations to Cooperative Actions
AI 导读
一项基于 Hanabi 衍生环境的研究测试了 8 个 LLM 在固定权重下接收决策的表现:线性探针能较准确解码发送方的意图约定,但模型的接收选择并不总与之一致。以通用规则形式给出约定信息带来的合作提升有限且因模型而异,而将其转化为动作建议平均收益更大;Qwen3-8B 案例中模型对动作建议的敏感度明显高于规则陈述。
正文
Abstract:Cooperation with unfamiliar partners requires adapting to communication conventions that are not known in advance. We study this problem in a controlled Hanabi-derived environment with scripted hint generation, LLM-controlled receiving decisions, and frozen model weights. Across eight LLMs, linear probes recover intent conventions substantially more accurately than target conventions, yet receiving choices do not consistently agree with the sender's convention. We compare probe-predicted and ground-truth conventions presented either as general rules or as externally computed action recommendations. Rule statements yield modest and model-dependent changes in cooperation, whereas action translation produces larger gains on average. In a Qwen3-8B case study, matched-state statement reversals reveal much greater sensitivity to action recommendations than to rule statements. Activation transfers from oracle-action and non-oracle hint-restatement donors improve intent accuracy on both action classes, but the tested alternatives do not reliably reproduce these benefits. Together, these results distinguish convention decodability, sensitivity to convention information, and cooperative performance, and highlight limitations in turning available partner information into receiving decisions.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.08129 [cs.LG] |
| (or arXiv:2610.08129v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08129 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yuhwan Jeong [view email]
[v1]
Tue, 6 Oct 2026 10:44:55 UTC (388 KB)
来源:arXiv:cs.LG · arxiv.org