跳到正文
arXiv:cs.CL· Akesh Gunathilake, Nadil Karunarathna, Tharusha Bandaranayake, Nisansa de Silva, Surangika Ranathunga, Nevidu Jayatilleke·· 3 小时前

LMSpell:用预训练语言模型做拼写纠错

LMSpell: Spell Correction with Pre-Trained Language Models

AI 导读

研究首次系统比较三类预训练语言模型在多语言(含低资源语言)拼写纠错上的效果,发现仅有 270M 参数的 Gemma 3 和 mBART50 在 5k 句数据上微调后即可胜过基于规则的纠错器。论文还以僧伽罗语为例,展示低资源语言拼写纠错的困境。

正文

View PDF HTML (experimental)

Abstract:Spell correction is still a challenging problem for many languages, especially low-resource languages (LRLs). While pre-trained language models (PLMs) have been employed for spell correction, there has been no proper comparison across PLMs. We present the first empirical study on the effectiveness of the three types of PLMs for spell correction across multiple languages, including low-resource languages. We show that even relatively small PLMs such as the 270M-parameter Gemma 3 and mBART50, when fine-tuned on a dataset of only 5k sentences, can outperform rule-based spell correctors, highlighting a practical pathway for building effective spell correction systems with limited data. We also present a case study with Sinhala to shed light on the plight of spell correction for LRLs.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2512.05414 [cs.CL]
  (or arXiv:2512.05414v4 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2512.05414

arXiv-issued DOI via DataCite

Related DOI: https://doi.org/10.1109/MERCon71835.2026.11691371

DOI(s) linking to related resources

Submission history

From: Akesh Samuditha Gunathilake [view email]
[v1] Fri, 5 Dec 2025 04:14:09 UTC (9,861 KB)
[v2] Mon, 8 Dec 2025 02:01:26 UTC (9,860 KB)
[v3] Thu, 11 Dec 2025 14:22:42 UTC (9,850 KB)
[v4] Thu, 8 Oct 2026 17:32:24 UTC (8,565 KB)

来源:arXiv:cs.CL · arxiv.org