arXiv:cs.LG· G. L. John Salvin (Indian Institute of Technology Palakkad), Swapnil Hingmire (Indian Institute of Technology Palakkad)·· 6 小时前AI 评分45
让 COMET 跨文字系统可比:Indic MT 评测中 tokeniser 引发的文字系统偏差诊断与修正
Making COMET Comparable Across Scripts: Diagnosis and Correction of Tokeniser-Induced Script Bias in Indic MT Evaluation
AI 导读
COMET 翻译质量分数依赖目标语言所用文字系统,在 IndicMT Eval 上将目标改写为拉丁字母后,文字系统身份可解释原生文字下 COMET 方差的 22.9%,所研究的五种语言与标注者一致性全部下降,根因指向 tokeniser。
正文
Abstract:COMET reports translation quality as a single number, and that number is routinely compared across target languages written in different scripts. Such a comparison assumes Script Invariance: the score should not depend on the writing system that carries the target. We test it on IndicMT Eval by re-encoding the target into Latin script, which changes orthographic form while holding content and human ratings fixed. Script identity then accounts for 22.9% of native-script COMET variance, and agreement with annotators falls in all five languages studied. We trace the effect to the tokeniser and measure it with three label-free diagnostics. The bias is two faults, not one. Scores from different scripts occupy incompatible ranges, and within a single script the metric orders translations less accurately. No order-preserving transform of the score can repair the second fault. The first is removed exactly by COMET-QN, which maps the score distribution of each (language, script) pair onto a shared reference. Pooled agreement with annotators rises from 0.300 to 0.399, which is what makes scores from different scripts safe to place on one axis, and every within-language ordering is provably preserved. A regressor over parity features recovers a further 17.1% of the lost sensitivity. The remainder belongs to the encoder, and no post-processing can reach it. We therefore recommend publishing the normalised score, the three diagnostics, and the identity of the tokeniser they were computed against, so that a reader can tell how much of a score reflects translation quality and how much reflects the writing system.
| Comments: | 18 pages, 2 figures. Camera-ready version, accepted at WMT 2026. Code and data: this https URL |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| ACM classes: | I.2.7 |
| Cite as: | arXiv:2610.08159 [cs.CL] |
| (or arXiv:2610.08159v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08159 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: G.L. John Salvin [view email]
[v1]
Tue, 6 Oct 2026 11:11:43 UTC (189 KB)
来源:arXiv:cs.LG · arxiv.org