跳到正文
arXiv:cs.AI· Jia Li, Yichao He, Yangchen Yu, Qiankun Li, Xinyi Li, Baiyi Ye, Zhenzhen Hu, Richang Hong, Erik Cambria·· 3 小时前

从表层到深度:面向多模态情感理解的认知评估推理

From Surface to Depth: Towards Cognitive Appraisal Reasoning in Multimodal Emotion Understanding

AI 导读

研究者提出将多模态情感理解从感知推进到认知评估的新范式,并发布 CogEmo-40K 指令微调数据集、紧凑稀疏 MLLM 模型 CogEmo-MoE 及评测基准 CogEmo-Bench。

正文

View PDF HTML (experimental)

Abstract:Recent multimodal large language models (MLLMs) increasingly incorporate explainable reasoning for emotion understanding. However, reasoning based mainly on observable affective cues can reduce emotion understanding to superficial cue-label associations, giving rise to the Clever Hans effect. Such shortcuts become unreliable when affective cues are implicit, conflicting across modalities, linguistically misleading, or obscured by redundant details. In contrast, human emotions are shaped by how individuals interpret and evaluate surrounding events beyond observable cues. Inspired by appraisal theories of emotion, we formulate multimodal emotion understanding as a progression from perception to cognitive appraisal, and introduce a dataset, a model, and a benchmark to support this novel paradigm. CogEmo-40K is a large-scale instruction-tuning dataset constructed through a perception-to-appraisal pipeline to elicit evidence-grounded reasoning across six cognitive appraisal dimensions underlying emotion. CogEmo-MoE is a compact sparse MLLM that introduces interleaved MoE blocks for appraisal-specific adaptation, enabling effective appraisal reasoning at a substantially smaller scale than typical emotion MLLMs. CogEmo-Bench introduces an Appraisal Evidence Quality Score (AEQS) to assess cognitive-affective understanding across six complementary appraisal dimensions, addressing the limitation of conventional emotion metrics that evaluate what emotion is predicted but not why it arises. Extensive experiments show that our paradigm not only leads CogEmo-Bench, but also exhibits strong cross-domain generalization. Our findings suggest that perception-to-appraisal reasoning can move beyond surface-level cue-label associations toward more reliable multimodal emotion understanding and closer cognitive alignment between MLLMs and humans.
Comments: 34 pages, 10 figures, Project page: this https URL
Subjects: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
Cite as: arXiv:2610.11918 [cs.MM]
  (or arXiv:2610.11918v1 [cs.MM] for this version)
  https://doi.org/10.48550/arXiv.2610.11918

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jia Li [view email]
[v1] Thu, 8 Oct 2026 13:14:39 UTC (3,715 KB)

来源:arXiv:cs.AI · arxiv.org