跳到正文
arXiv:cs.CL· Parisa Suchdev, Juniper Lovato·· 3 小时前AI 评分54

审计 Claude Opus 4、GPT-4.1 与 Gemini 2.5 Pro 跨语言改编数学应用题的文化翻译质量

Who Brought Easter Eggs to Eid? Auditing LLM-Generated Cultural Translation of Math Word Problems Across Languages and Regions

AI 导读

研究者审计 Claude Opus 4、GPT-4.1 和 Gemini 2.5 Pro 将 60 道英文数学应用题改编为孟加拉语、印地语、旁遮普语、乌尔都语、信德语、意大利语和西西里语的表现,标注 6489 条实体转换。

正文

View PDF HTML (experimental)

Abstract:Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether those adaptations are consistent across models, preserve cultural diversity at scale, and reveal which cultural entities models treat as most salient. We analyze how Claude Opus 4, GPT-4.1, and Gemini 2.5 Pro adapt 60 English math word problems into Bengali, Hindi, Punjabi (India), Urdu, Sindhi (Pakistan), Italian, and Sicilian (Italy), a language set spanning the full resource spectrum, from high-resource Italian and Hindi to under-studied Sindhi, Sicilian, and Punjabi. We annotate 6,489 entity transformations, coding whether models preserve, localize, generalize, omit, or change entities such as names, foods, and places. Models agree on transformation type in 62.5% of cases and on specific substitutions in only 33.5%, meaning model choice directly shapes which cultural world students encounter. All 21 language-model combinations show entropy collapse, with adaptation compressing rather than expanding cultural diversity. Models prioritize surface markers such as names, foods, and currencies while preserving deeper structural features such as grade-level systems that embed culturally specific assumptions. Despite prompts specifying target countries, models misattribute regional context by using Bangladeshi taka for Indian Bengali students and produce cross-cultural contamination, such as adapting egg hunts as Eid activities. Some failures are visible in individual translations. Others, including diversity collapse, systematic preference for surface markers, and consistent regional misattribution, emerge only through corpus-level analysis. The surface plausibility that makes adapted problems look correct is precisely what makes deeper failures easy to overlook.
Comments: 18 pages total with references and appendix, 9 figures, accepted at AIES
Subjects: Computation and Language (cs.CL); Computers and Society (cs.CY)
Cite as: arXiv:2606.11009 [cs.CL]
  (or arXiv:2606.11009v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2606.11009

arXiv-issued DOI via DataCite

Submission history

From: Parisa Suchdev [view email]
[v1] Tue, 9 Jun 2026 15:50:12 UTC (321 KB)
[v2] Wed, 7 Oct 2026 15:30:02 UTC (332 KB)

来源:arXiv:cs.CL · arxiv.org