跳到正文
arXiv:cs.CL· Chaymaa Abbas, Nour Shammaa, Mariette Awad·· 3 小时前

翻译如何掩盖数据污染:来自阿拉伯语语料的证据

Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora

AI 导读

研究将 MMLU 和 XQuAD 评测题目翻译成阿拉伯语,在递增曝光水平下喂给四个开源指令微调 LLM,再回到英文原题评测。英文专用检测方法 TS-Guessing 和 Min-K%++ 的信号在翻译曝光下基本消失,Min-K%++ 处于随机水平或更低,但英文 MMLU 成绩仍随阿拉伯语曝光上升。

正文

View PDF HTML (experimental)

Abstract:Data contamination can invalidate benchmark evaluation when a model benefits from memorized evaluation content rather than genuine generalization. Yet contamination is difficult to audit when the exposed content differs in language from the evaluation benchmark. We study this failure mode by deliberately exposing four open-weight instruction-tuned LLMs to Arabic translations of MMLU and XQuAD evaluation items at increasing exposure levels, then evaluating them on the original English tasks. This controlled setup is a proxy for contamination rather than a reconstruction of real-world pretraining leakage. We first test two English-centric post-hoc probes, TS-Guessing and Min-K%++, and find that their signals largely disappear under translated exposure: TS-Guessing remains weak except for model-specific positional recall on MMLU, while Min-K%++ stays at or below chance. At the same time, English MMLU performance increases with Arabic exposure, showing that the absence of an English contamination signal does not imply the absence of an exposure effect. We then introduce Translation-Aware Contamination Detection (TACD), a training-data-free diagnostic based on cross-lingual prediction consistency and choice reordering. Cross-lingual consistency is substantially higher than an independence baseline and generally increases relative to the clean condition, although its magnitude is model-dependent and not strictly monotonic. These results show that translation can conceal contamination-related effects from English-only probes and motivate multilingual diagnostics that are explicitly framed as evidence of contamination-consistent behavior rather than definitive membership tests.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2601.14994 [cs.CL]
  (or arXiv:2601.14994v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2601.14994

arXiv-issued DOI via DataCite

Submission history

From: Chaymaa Abbas [view email]
[v1] Wed, 21 Jan 2026 13:53:04 UTC (2,134 KB)
[v2] Thu, 8 Oct 2026 14:34:01 UTC (3,779 KB)

来源:arXiv:cs.CL · arxiv.org