arXiv:cs.CL· Hala Almaghout, Christian Federmann, Qin Gao·· 3 小时前
大语言模型用于机器翻译质量标注:人类与模型都面临挑战
Large Language Models for Machine Translation Quality Annotation: Humans and Models Are Both Challenged
AI 导读
研究评估了 LLM 在 MQM 和 ESA 两种机器翻译质量评估方案下与人类标注者的一致性,覆盖 70 个语言对的长上下文测试集及 WMT23、WMT25 公开数据。
正文
Abstract:Large Language Models (LLMs) are considered to be a more efficient and cost-effective alternative to human judgment for Machine Translation (MT) evaluation. With MT evaluation spanning a large number of language pairs, domains and levels of annotation granularity, LLMs must be thoroughly evaluated across these dimensions before being reliably used as alternatives to human evaluation. In this paper, we evaluate the performance of LLMs for two prominent MT quality evaluation schemes: Multidimensional Quality Metrics (MQM) and Error Span Annotation (ESA) by comparing their agreement with human annotators. We present results on a long-context test set of 70 language pairs and the publicly available WMT23 and WMT25 data, investigating both score and error span annotation agreement across a variety of language pairs and domains. Our results show that while LLM agreement with human annotators exceeds agreement between human annotators for some evaluation tasks, both vary substantially across annotation schemes, language pairs and domains and remain unreliable for most settings. Furthermore, we identify challenges facing both human and LLM annotators: humans are particularly challenged by fine-grained MQM annotations and low-resource language pairs, while LLMs struggle with minor errors, wrong language variants and error span annotation. Our results highlight the potential for improvement for both human and LLM annotation performance, possibly through human-LLM collaborative annotation pipelines that address the reliability issues identified in this work.
| Comments: | Accepted at EMNLP 2026 |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.10918 [cs.CL] |
| (or arXiv:2610.10918v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10918 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hala Almaghout [view email]
[v1]
Wed, 7 Oct 2026 21:17:26 UTC (1,024 KB)
来源:arXiv:cs.CL · arxiv.org