跳到正文
arXiv:cs.CL· Himarsha R. Jayanetti, Sivakanesan Dhanushkanda, Shuai Hao, Michael L. Nelson, Michele C. Weigle·· 3 小时前AI 评分34

研究:人类、专用工具与 LLM 在社交媒体文本情感分析上难达一致

Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts

AI 导读

一项针对 100 条推文的研究对比 TextBlob、VADER、Twitter-roBERTa-base 与 Qwen3-32B、GPT-OSS-120B、Llama-4-Maverick-17B 及 6 名人类标注者的情感判断,发现即便人类之间也仅达一般一致性。

正文

View PDF HTML (experimental)

Abstract:Social media is a rich source of real-time public sentiment, but widely used sentiment analysis tools are often applied without understanding their limitations. In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets. We measured agreement using two statistical measures: Cohen's kappa for pairwise comparisons and Fleiss' kappa for multiple raters. Even among the human raters, our results showed only fair agreement, highlighting the subjectivity of sentiment analysis. Higher agreement was observed under the binary sentiment classification (negative vs. non-negative and positive vs. non-positive) than under the three-class classification across both humans and automated tools. The Twitter-roBERTa-base model showed the strongest alignment with human ratings, outperforming both bespoke sentiment tools and LLMs, particularly in distinguishing negative versus non-negative sentiment. LLMs showed substantial agreement among themselves and moderate to substantial alignment with humans, performing better in positive vs. non-positive classifications. Our findings underscore that domain-specific fine-tuning remains crucial for reliable social media sentiment analysis, and human-centered evaluation remains essential for establishing gold-standard labels.
Comments: 11 pages, 1 figure, 2 tables, accepted for publication at TPDL 2026
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.10318 [cs.CL]
  (or arXiv:2610.10318v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.10318

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Himarsha R Jayanetti [view email]
[v1] Wed, 7 Oct 2026 16:11:12 UTC (247 KB)

来源:arXiv:cs.CL · arxiv.org