arXiv:cs.CL· Tobias Hallmen, Fabian Deuser, Robin-Nico Kampa, Norbert Oswald, Elisabeth Andr\'e·· 7 小时前AI 评分47
细粒度情绪识别基准的失败在于读取方式而非感知能力:EmoNet-Face-HQ 评测方法被指有误
The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception
AI 导读
针对 EmoNet-Face-HQ 细粒度情绪识别基准,研究者发现现成 VLM 改用逐类别二分类 logits 读取答案后,表现可匹敌甚至超越专用微调模型 Empathic-Insight-Face(EIF Small/Large)。
正文
Abstract:Fine-grained emotion recognition supports therapy tools and social robots, but it needs facial data, which raises privacy and data-protection concerns. EmoNet-Face-HQ answers that with generated portraits, expert-rated over a $40$-category taxonomy far finer than the usual six to eight basic emotions. Under the protocol it ships with, vision-language models (VLMs) score poorly on that taxonomy, and the benchmark concludes that a dedicated fine-tuned model is necessary: Empathic-Insight-Face (EIF; Small/Large). We show that off-the-shelf VLMs match or beat that fine-tuned model when the answer is not generated but read from the logits, as one binary query per category. We keep the benchmark's images, taxonomy and ratings, and change only how the answer is read. Experts agree at $\kappa_w = 0.468$ on the five categories they measure most reliably. Generatively, no interval among eleven open-weight VLMs lies entirely above that anchor ($\kappa_w=0.268$-$0.486$). Under verification all eleven clear it, each of them significantly better at $\kappa_w=0.507$-$0.586$. Three also significantly beat EIF sitting at $\kappa_w = 0.551$ (Small; $0.534$ Large). The gain comes from the graded probability and not from asking a yes/no question: as a control, thresholding those same probabilities to yes/no costs 142% of the average gains and drops binarization below generative elicitation to $\kappa_w=0.254$-$0.423$. A replication on real photographs (FACES) is weaker and mixed: of the ten models that pass a validity gate, six gain, three are neutral to positive and one is negative, so the effect is not confined to synthetic data.
| Comments: | Preprint. 19 pages, 6 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.08162 [cs.CV] |
| (or arXiv:2610.08162v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08162 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Fabian Deuser [view email]
[v1]
Tue, 6 Oct 2026 11:14:23 UTC (106 KB)
来源:arXiv:cs.CL · arxiv.org