arXiv:cs.AI· Sree Bhattacharyya, Evgenii Kuriabov, Lucas Craig, Tharun Dilliraj, Reginald B. Adams, Jr., Jia Li, James Z. Wang·· 7 小时前AI 评分43
LLM 与人类的情感概念表征差异研究:模型更分类化、同质化
Too Categorical to be Human: Emotion Concepts in LLMs and Humans
AI 导读
研究者提出用"行为表征"刻画情感概念,基于认知评价理论构建了覆盖 15 种情绪类别的情绪场景基准数据集,对比 LLM 与人类的情感概念结构相似性。结果显示 LLM 的情感表征比人类更分类化、同质化和确定化,单一情绪概念的内部多样性更低,不同情绪间的距离更远,且这种结构对任务框架和人口学角色等语境变化保持稳健。模型检查点分析表明,表征的离散化特征出现在训练中期阶段之后,且不受不同后训练策略影响。
正文
Abstract:Understanding human emotions is central to user-facing AI applications, safety alignment, and the simulation of human behavior. As emotional stimuli shape high-stakes behavior in Large Language Models (LLMs), there is increasing interest in how models represent emotion concepts internally. Mechanistic accounts of these representations, however, cannot be compared directly against humans: emotion processing in humans is highly distributed and yields no equivalent neural representation. To understand whether LLMs internalize emotion concepts in a way similar to humans, we propose characterizing the abstract concept of an emotion using external behavioral signatures, which we term behavioral representations. Using the theory of cognitive appraisals, which enables representing emotional situations along interpretable evaluative dimensions, we create a benchmark dataset of emotional scenarios spanning 15 emotion categories. We elicit behavioral representations of emotion concepts from LLMs and humans using our benchmark, and study their structural similarity. We find that LLMs represent emotion concepts more categorically, homogeneously, and determinately than humans, representing a single emotion concept with less internal diversity, and place different emotions further apart. The categorical structure of representations in LLMs is further robust to contextual variation, including with different task framing and demographic personas. Analyzing model checkpoints across different training stages, we also find that the discretized nature of representations appears after the mid-training stage itself and is unaffected by different post-training strategies. Through our results, we highlight a key difference in how LLMs behaviorally represent emotion concepts, curbing the subjectivity inherent to the human experience of emotions.
| Comments: | 19 pages of main body; A version was presented at WiML Workshop @ NeurIPS 2025 |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2508.05880 [cs.CL] |
| (or arXiv:2508.05880v3 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2508.05880 arXiv-issued DOI via DataCite |
Submission history
From: Sree Bhattacharyya [view email]
[v1]
Thu, 7 Aug 2025 22:19:15 UTC (1,156 KB)
[v2]
Fri, 13 Mar 2026 17:27:05 UTC (5,870 KB)
[v3]
Tue, 6 Oct 2026 01:35:04 UTC (797 KB)
来源:arXiv:cs.AI · arxiv.org