跳到正文
arXiv:cs.LG· G. L. John Salvin (Indian Institute of Technology Palakkad), Swapnil Hingmire (Indian Institute of Technology Palakkad)·· 6 小时前AI 评分45

让 COMET 跨文字系统可比:Indic MT 评测中 tokeniser 引发的文字系统偏差诊断与修正

Making COMET Comparable Across Scripts: Diagnosis and Correction of Tokeniser-Induced Script Bias in Indic MT Evaluation

AI 导读

COMET 翻译质量分数依赖目标语言所用文字系统,在 IndicMT Eval 上将目标改写为拉丁字母后,文字系统身份可解释原生文字下 COMET 方差的 22.9%,所研究的五种语言与标注者一致性全部下降,根因指向 tokeniser。

正文

View PDF HTML (experimental)

Abstract:COMET reports translation quality as a single number, and that number is routinely compared across target languages written in different scripts. Such a comparison assumes Script Invariance: the score should not depend on the writing system that carries the target. We test it on IndicMT Eval by re-encoding the target into Latin script, which changes orthographic form while holding content and human ratings fixed. Script identity then accounts for 22.9% of native-script COMET variance, and agreement with annotators falls in all five languages studied. We trace the effect to the tokeniser and measure it with three label-free diagnostics. The bias is two faults, not one. Scores from different scripts occupy incompatible ranges, and within a single script the metric orders translations less accurately. No order-preserving transform of the score can repair the second fault. The first is removed exactly by COMET-QN, which maps the score distribution of each (language, script) pair onto a shared reference. Pooled agreement with annotators rises from 0.300 to 0.399, which is what makes scores from different scripts safe to place on one axis, and every within-language ordering is provably preserved. A regressor over parity features recovers a further 17.1% of the lost sensitivity. The remainder belongs to the encoder, and no post-processing can reach it. We therefore recommend publishing the normalised score, the three diagnostics, and the identity of the tokeniser they were computed against, so that a reader can tell how much of a score reflects translation quality and how much reflects the writing system.
Comments: 18 pages, 2 figures. Camera-ready version, accepted at WMT 2026. Code and data: this https URL
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
ACM classes: I.2.7
Cite as: arXiv:2610.08159 [cs.CL]
  (or arXiv:2610.08159v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08159

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: G.L. John Salvin [view email]
[v1] Tue, 6 Oct 2026 11:11:43 UTC (189 KB)

来源:arXiv:cs.LG · arxiv.org